Antonio OrvietoView profile
Lecturer
Antonio Orvieto is a Lecturer at the University of Tübingen, Principal Investigator at the ELLIS Institute Tübingen, and independent group leader at the Max Planck Institute for Intelligent Systems. He leads the Deep Models and Optimization research group and serves as faculty for the CLS, ELLIS, and IMPRS-IS PhD Programs. Orvieto holds a PhD from ETH Zürich and has conducted research at DeepMind London, Meta (FAIR) Seattle, MILA, and Inria Paris. Orvieto's research focuses on improving the efficiency and accessibility of deep learning technologies through theoretical advancements in optimization and architecture design. His work spans two main areas: understanding large-scale optimization dynamics and designing innovative neural network architectures capable of reasoning with complex sequential data. His research has significant implications across multiple domains including biology, neuroscience, natural language processing, and music generation. His approach combines rigorous theoretical analysis with practical applications to address fundamental challenges in deep learning. His recent publications reveal a strong emphasis on understanding optimization landscapes, developing efficient recurrent architectures, analyzing transformer behavior, and exploring the theoretical foundations of state-space models. The work shows a consistent pattern of bridging theoretical insights with practical implementations, particularly in handling sequential data and improving training efficiency. Schmidt Sciences AI2050 Early Career Fellow Orvieto actively mentors PhD students and researchers in his Deep Models and Optimization group, which includes PhD candidates working on various aspects of deep learning theory and applications. His research is supported through his positions at the University of Tübingen, ELLIS Institute, and Max Planck Institute for Intelligent Systems. He has collaborated with leading researchers across multiple institutions including ETH Zürich, DeepMind, Meta, MILA, and Inria. His research group focuses on investigating the interplay between optimizers and architectures in deep learning, with particular emphasis on developing new networks for long-range reasoning. The group strongly believes that deep learning will revolutionize science and technology, and they aim to provide theoretical foundations that will enable scientists and engineers with limited resources to leverage powerful deep learning solutions.










