
معرفی
Jacob Steinhardt is an Assistant Professor in the Department of Statistics at UC Berkeley, where he is also part of BAIR (Berkeley Artificial Intelligence Research) and CLIMB. His research focuses on ensuring machine learning systems are understood by and aligned with humans, addressing critical challenges in AI safety and reliability.
Dr. Steinhardt's research centers on three main directions: Robustness - developing models resilient to distributional shifts, adversaries, and model mis-specification; Reward specification and reward hacking - creating methods to infer complex value functions from data and prevent degenerate policies; and Scalable alignment - designing ML systems that conform to interpretable abstractions despite their large scale. His work rethinks both theoretical and empirical paradigms of machine learning to address these critical challenges to AI safety.
An analysis of his recent publications reveals a consistent focus on understanding the internal mechanisms of neural networks, particularly large language models, to improve their alignment with human values. His research spans interpretability techniques, safety mechanisms, and evaluation frameworks, with significant contributions to understanding reward hacking, distributional shift, and model transparency. His work often combines theoretical insights with empirical validation through novel experimental frameworks.
Dr. Steinhardt actively mentors a diverse group of PhD students including Ruiqi Zhong (co-advised with Dan Klein), Meena Jagadeesan (co-advised with Mike Jordan), Erik Jones (co-advised with Anca Dragan), and several others. His former students have gone on to positions at leading organizations including OpenAI, Genentech, and the Center for AI Safety.
He is the Founder & CEO of Transluce, a non-profit research lab building open, scalable technology for understanding frontier AI systems. Through this initiative and his academic work, he contributes significantly to the growing field of AI safety research, bridging theoretical foundations with practical applications to make machine learning systems more reliable and beneficial.
Jacob Steinhardt در سایتهای دیگر
جستوجوهای مرتبط
شاید اینها هم برایتان مناسب باشند
Anca DraganUniversity of California, Berkeley · دانشیار
Anca DraganCarnegie Mellon University · دانشیار- FF. RussellUniversity of California, Berkeley · استاد
Alexander PanCalifornia Institute of Technology (Caltech) · پژوهشگر- DDaniel S. BrownUniversity of Utah · استادیار
- SStuart RussellUniversity of Trier · استاد