Andrea Vedaldi is a Professor of Computer Vision and Machine Learning at the University of Oxford's Department of Engineering Science, affiliated with the Visual Geometry Group (VGG). He specializes in unsupervised methods for understanding images and videos, focusing on 3D geometry and semantics. His research bridges foundational AI and practical applications, with contributions to generative models, neural fields, and self-supervised learning. Education: PhD in Computer Science (2008), University of California, Los Angeles MSc in Computer Science (2005), UCLA BSc in Information Engineering (2003), University of Padua Research Interests: Unsupervised learning, 3D perception, generative AI, neural rendering, and scalable vision systems. His work emphasizes ethical, responsible AI aligned with ERC-funded projects like UNION (ERC Consolidator Grant). Key Contributions: Co-developer of VLFeat and MatConvNet libraries Leader in 3D reconstruction and diffusion models (e.g., CatFree3D) Recipient of the PAMI Thomas S. Huang Prize and multiple best paper awards Grants & Service: Principal Investigator on £2.3M ERC Consolidator Grant (UNION) Co-organizer of major conferences (ECCV 2020 Program Chair, CVPR 2023 Area Chair) Reviewer for top journals/conferences (PAMI, CVPR, NeurIPS) Labs & Teams: VGG Group at Oxford, collaborating on projects like Meta 3D Gen and Common Objects in 3D (CO3D).
Deva Kannan Ramanan is a Professor at the Robotics Institute of Carnegie Mellon University , focusing on computer vision , machine learning , and human-centered robotics . His work bridges neurorobotics and visual perception , with applications in autonomous driving and 4D reconstruction . Research Topics Computer Vision 3-D Vision and Recognition Visual Servoing Neurorobotics Human-Centered Robotics Graphics & Creative Tools His recent publications in CVPR , ICRA , and ICCV emphasize 4D human reconstruction , neural rendering , and vision-language models for autonomous systems. He serves as General Chair of CVPR 2027 and Program Chair of CVPR 2018 , with IARPA funding for aerial-ground rendering (2023-2027). Current students include PhD candidates Sally Chen, Kangle Deng, and Zhiqiu Lin, while past advisees like Arun Vasudevan and Olga Russakovsky now hold positions at Amazon and Meta respectively.
Kai-Wei Chang is an Associate Professor at the University of California, Los Angeles (UCLA) in the Department of Computer Science, part of the Henry Samueli School of Engineering. He is also an Amazon Scholar at Alexa AI, focusing on advancing trustworthy AI and multimodal foundation models. His research bridges NLP, machine learning, and ethical AI, with a focus on fairness, robustness, and bias mitigation in language and vision-language systems. Education: Ph.D. in Computer Science (UIUC, 2015), M.S. and B.S. in Computer Science and Electrical Engineering from National Taiwan University. Research Interests: Trustworthy NLP (fairness, robustness), Multimodal Foundation Models (e.g., VisualBERT, GLIP), Reasoning in LLMs, and mitigating societal biases in AI systems. Notable contributions include pioneering work on aligning NLP models with human values and developing SOTA multimodal models like DesCo and GLIP. Awards: Sloan Research Fellowship (2021), Okawa Grant (2018), EMNLP Best Paper (2017), KDD Best Paper (2010). His work is funded by NSF, IARPA, ONR, and industry partners like Amazon, Google, and Facebook. Service: VP-Elect of SIGDAT, Ethics Committee Chair (NAACL 2022), Organizer of Trustworthy NLP Workshops, and Senior Area Chair for top conferences (ACL, NeurIPS, AAAI). Labs/Teams: Leads the UCLA Natural Language Processing Group, fostering interdisciplinary research on ethical AI and multimodal systems.
Christian Theobalt is a Professor of Computer Science at Saarland University and Scientific Director of the Visual Computing and Artificial Intelligence Department at the Max Planck Institute for Informatics . He leads the Saarbruecken Center for Visual Computing as a strategic partnership between Google and MPI. PhD in Computer Science (2005) from MPI-INF/Saarland University Postdoctoral Researcher at MPI (2005-2007) Visiting Assistant Professor at Stanford (2007-2009) His research focuses on the intersection of Computer Graphics, Computer Vision, and Artificial Intelligence , with specializations in: 3D/4D Human Reconstruction Neural Rendering Performance Capture Geometric Deep Learning Volumetric Video Quantum Visual Computing Recent article trends show emphasis on: Quantum computing applications in visual reconstruction Neural rendering with radiance fields Human motion capture from egocentric views Gesture-language interaction modeling Multi-modal scene understanding Real-time free-viewpoint rendering Scientific Awards : Fellow of EUROGRAPHICS (2022) CVPR Best Student Paper Honorable Mention (2020) ERC Consolidator Grant (2017) Karl Heinz Beckurts Award (2017) Busy Beaver Teaching Award (2016) ERC Starting Grant (2013) Advising over 50 students and researchers including: Current researchers: Viktor Rudnev, Linjie Lyu, Mohit Mendiratta Postdocs: Kwang In Kim, Kiran Varanasi Alumni: Franziska Mueller (Google), Dushyant Mehta (Qualcomm), Ayush Tewari (MIT) Grants include ERC grants, Google Glass Research Award, and multiple industry partnerships. His lab maintains cutting-edge facilities with: Multi-camera capture systems Quantum annealing infrastructure HDR radiance field technology GPU clusters for AI research Time-of-flight imaging systems Event camera arrays
Amir Zamir is a tenure-track Assistant Professor of Computer Science at the Swiss Federal Institute of Technology Lausanne (EPFL) in the School of Computer & Communication Sciences. Previously, he worked at UC Berkeley, Stanford, and UCF with prominent researchers including Silvio Savarese, Jitendra Malik, Mubarak Shah, Rahul Sukthankar, and Leonidas Guibas. He currently leads the Visual Intelligence & Learning Lab at EPFL and serves as chief scientist of Duranta, having previously been the CVML chief scientist of Aurora Solar (a Forbes AI 50 company valued at $4B in 2022) from 2015 to 2022. Dr. Zamir's research spans computer vision, machine learning, and artificial intelligence, with a focus on developing general multi-modal/multi-task vision systems that operate as active agents in the real world. His work emphasizes slow science principles, seeking fundamental understanding over quick publications. Key research areas include embodied vision, multimodal foundation models, computational imaging, and vision-language integration. His notable projects include 4M, Taskonomy, Gibson Environment, Omnidata, and MultiMAE, which have significantly influenced the field of computer vision and embodied AI. Zamir has made substantial contributions to the computer vision community through numerous high-impact publications and leadership roles. His work demonstrates a consistent focus on creating vision systems that go beyond narrow and passive methods toward more general, active, and embodied approaches. The trajectory of his research shows increasing sophistication in handling multiple modalities and tasks within unified frameworks, culminating in recent work on multimodal foundation models that can handle diverse vision tasks. Dr. Zamir has received numerous prestigious awards including the Young Researcher Award 2022 from ECCV, the PAMI Mark Everingham Prize 2022, SIGGRAPH 2022 Best Paper Award for CLIPasso, CVPR 2018 Best Paper Award for Taskonomy, and CVPR 2016 Best Student Paper Award. He is also an ELLIS Faculty Scholar and received the NVIDIA Pioneering Research Award in 2018 for the Gibson Environment. As an advisor, Dr. Zamir has mentored numerous PhD students including Roman Bachmann, Andrei Atanov, Rishubh Singh, Jason Toskov, Kunal Pratap Singh, Zhitong Gao, Mingqiao Ye, and Muhammad Uzair Khattak. His former PhD students include Oguzhan Kar (now at Apple), Alexander Sasha Sax (co-advised with Jitendra Malik, now at Meta FAIR), and Teresa Yeo (now at MIT-Singapore Alliance). Dr. Zamir teaches several advanced courses including CS-503 Visual Intelligence, CS-500 AI Product Management, COM-304 Intelligent Systems, and ENG-615 Topics in Autonomous Robotics. Dr. Zamir leads the Visual Intelligence & Learning Lab at EPFL, which focuses on developing fundamental methods for visual intelligence that can operate effectively in real-world environments. The lab takes an interdisciplinary approach combining computer vision, machine learning, robotics, and cognitive science to create systems that can perceive, understand, and interact with the world. Current research directions include multimodal foundation models, computational imaging, embodied vision, and personalization of generative models.
Shoudong Huang is a Professor at the School of Mechanical and Mechatronic Engineering , University of Technology Sydney, and Deputy Director of the UTS Robotics Institute. His research focuses on mobile robot navigation , SLAM , nonlinear state estimation , and surgical robotics . He has published over 200 papers and is recognized as one of the 100 Most Influential Scholars in Robotics (Aminer, 2018). PhD in Automatic Control, Northeastern University (China) Postdoctoral Research Fellow, University of Hong Kong (1998-2000) Research Fellow, Australian National University (2001-2003) Full-time academic roles at UTS since 2004 His work addresses challenges in robot localization across extreme environments (underwater, underground mining, surgical settings) and develops globally optimal SLAM algorithms with guaranteed performance. He has secured over $4 million AUD in external funding, including ARC Discovery grants and industry partnerships. Recent publications emphasize cross-modal calibration (camera-LiDAR), interval analysis for bounded noise , and template-based deformable surface reconstruction . These span applications in autonomous driving, surgical navigation, and UAV guidance. Chancellor’s Medal for Research Excellence (2020) Supervisor of the Year (2023) Best Paper Award (2016 ICARCV) Huang serves as Associate Editor for IEEE Transactions on Robotics and International Journal of Robotics Research , and has held leadership roles in top robotics conferences like IROS and RSS. His collaborations span MIT, USC, Zhejiang University, and industry partners including PMSW Research Pty Ltd and Multiplex Constructions Pty Ltd.
Vineeth N Balasubramanian is a Professor in the Department of Computer Science & Engineering at the Indian Institute of Technology Hyderabad, with affiliate faculty status in the Department of Artificial Intelligence. His research focuses on the intersection of deep learning, machine learning, and computer vision, emphasizing explainability, robustness, and real-world applications. He leads Lab 1055, which investigates problems such as Explainable and robust AI/ML systems Lifelong learning in evolving environments Multimodal vision-language models Applications in agriculture, autonomous navigation, and human behavior analysis His recent work includes causal reasoning in transformers, vision-language model capabilities, and drone-based object detection. Funded by organizations like Google, Microsoft, Intel, and DST, he has received multiple awards including the World's Top 2% Scientists (2022-23), INSA/INAE Fellowships, and Best Paper recognitions. Lab 1055 collaborates with institutions like CMU, UBC, and Monash University, contributing to cutting-edge advancements in AI.
Xingxing Zuo is an Assistant Professor (tenure-track) in the Robotics Department at MBZUAI. He holds a PhD from Zhejiang University (2021) and a Bachelor’s from UESTC (2016). Previously, he was a Postdoctoral Scholar at Caltech (2024–2025), a Postdoc at ETH Zurich (2019–2021), and held visiting roles at TU Munich, University of Delaware, and University of Technology Sydney. His research focuses on robotics, 3D computer vision, and embodied AI, with emphasis on robot-human collaboration, state estimation, and sensor fusion. Educations: PhD in Robotics, Zhejiang University (2021, with honors) Bachelor’s in Computer Science, University of Electronic Science and Technology of China (2016, with honors) Research Highlights: Develops novel methods for LiDAR-camera-inertial fusion, neural radiance fields, and radar-cameras systems Pioneered techniques like Flying Co-Stereo (long-range aerial mapping) and FMGS (vision-language embedded 3D splatting) Focuses on real-time SLAM, robust depth estimation, and photorealistic scene reconstruction Awards & Recognition: Best Paper Finalist at ICRA 2021 (CodeVIO) Oral Presentation at ICCV 2021 (MBA-VO) Recipient of Google Visiting Faculty Researcher (2023) Grants & Labs: Organized Thermal Infrared in Robotics workshop at ICRA 2025 Leads research on embodied AI and multi-sensor SLAM systems Develops open-source tools like LIC-Fusion and Coco-LIC frameworks
Adriana Kovashka is an Associate Professor at the University of Pittsburgh , affiliated with the School of Computing and Information and serving as Department Chair . Her academic journey began with BA degrees in Computer Science and Media Studies from Pomona College (2008) and a PhD in Computer Science from The University of Texas at Austin (2014). Joined Pitt’s faculty in January 2015 NSF CAREER awardee (2021) Google Faculty Research Award recipient Dr. Kovashka’s research spans Computer Vision , Machine Learning , and Natural Language Processing , focusing on visual rhetoric, weak multimodal supervision, and domain adaptation. She pioneered techniques for analyzing political imagery, developing robust object detection frameworks, and exploring the intersection of visual and textual persuasion through large-scale annotated datasets. Her recent work emphasizes geographic diversity in vision-language systems, audio-visual fusion for domain generalization, and shape-texture bias mitigation in CNNs. Key publications include groundbreaking studies on symbolic reasoning, multimodal dialogue systems, and ethical AI applications in education. Scientific honors include: NSF CRII Award (2016) NSF CAREER Award (2021) Pitt CRDF Award (2016, 2018) Best Paper at ECV Workshop (2021) Google Faculty Research Award (2016, 2018) Dr. Kovashka actively mentors students in multimodal learning projects and collaborates with interdisciplinary teams on NSF-funded initiatives. She co-organizes workshops like the first CVPR workshop on advertisement understanding and leads research groups exploring human-AI co-learning systems.
Katerina Fragkiadaki is the JPMorgan Chase Associate Professor of Computer Science in the Machine Learning Department at Carnegie Mellon University. She works at the intersection of Artificial Intelligence, Computer Vision, Machine Learning, Language Understanding, and Robotics. PhD from GRASP Lab, University of Pennsylvania Postdoctoral researcher at UC Berkeley (with Jitendra Malik) and Google Research Recipient of NSF CAREER, DARPA Young Investigator, Amazon, Google, Sony, UPMC, and AFOSR awards Organizer of CoRL 2023 Workshop on Generalist Robots ICLR 2024 Program Chair, multiple area chair roles Her research group focuses on developing machines that autonomously improve world models through human-environment interactions, with specific emphasis on: Representation learning and video understanding 2D/3D unified vision-language models Generative simulation and reinforcement learning Real2Sim/Sim2Real robot learning Continual learning and spatial common sense 3D scene reconstruction and dynamics Recent publications highlight advancements in: 3D mesh generation with compositional transformers Unified 2D/3D perception frameworks Physics-aware generative models Diffusion-based robotic manipulation policies Embodied agents with memory prompting Awards include: 2024: DARPA Young Investigator Award 2023: Amazon Faculty Award 2022: Sony Faculty Research Award 2021: UPMC Faculty Research Award 2020: NSF CAREER Award 2019: Google Faculty Award Key collaborations span institutions including UC Berkeley, Google Research, Stanford, MIT, and University of Tsukuba. Her work bridges theoretical innovation with practical applications in: Autonomous robot manipulation 4D world modeling Language-grounded perception Visual dynamics prediction Embodied program synthesis Physics-based simulation engines
Robin Jia is an Assistant Professor in the Thomas Lord Department of Computer Science at the University of Southern California (USC) , where he leads the AI, Language, Learning, Generalization, and Robustness (Allegro) Lab . His research focuses on enhancing the reliability and robustness of large language models (LLMs) through mechanistic understanding, benchmarking under distribution shifts, and neurosymbolic integration. Key affiliations include collaborations with the USC Keck School of Medicine and contributions to legal frameworks like the EU's Digital Services Act. Robin's research spans multiple domains, including: Scientific analysis of LLM capabilities in in-context learning , data memorization , and numerical reasoning Advancements in robust NLP systems , emphasizing uncertainty estimation and calibration Development of methods combining LLMs with symbolic solvers for complex reasoning tasks Interdisciplinary applications in medicine and law , such as privacy-preserving synthetic data generation and medical misconception evaluation . His recent publications (2024-2025) address Fourier-based numerical embeddings (NeurIPS), neurosymbolic planning (NAACL), and multimodal benchmarking (COLM), with a strong emphasis on privacy , fairness , and transparency . Scientific awards include the Google Research Scholar Award (2023) , SoCalNLP Symposium Best Paper Awards , and ACL/EMNLP outstanding papers . He advises PhD students like Johnny Wei and Ameya Godbole, and has secured grants from the NSF , USC-Capital One , and USC-Amazon .
Alexander Schwing is an Associate Professor in the Department of Electrical and Computer Engineering and Computer Science at the University of Illinois at Urbana-Champaign, affiliated with the Coordinated Science Laboratory. His research focuses on machine learning and computer vision with applications in 3D scene understanding, generative modeling, and multi-agent systems. Education: Diploma in Electrical Engineering and Information Technology, Technical University of Munich (TUM) PhD in Computer Science, ETH Zurich Postdoctoral Fellow, University of Toronto Research Interests: Structured prediction in deep learning Generative adversarial networks and stability Multi-modal vision-language models 3D scene reconstruction from single images Embodied agent collaboration Semantic segmentation with temporal coherence Recent Publications: Highlight trends in neural rendering, video object segmentation, and reinforcement learning with applications to 3D modeling and multi-agent systems. Notable innovations include SAIL-VOS dataset for amodal segmentation and NeRFDeformer for single-view scene transformation. Scientific Awards: NSF CAREER Award, 3M and Amazon research awards, multiple student recognition awards, ETH Zurich PhD medal, and best paper at Intelligent Tutoring Systems 2014. Teaching: Offers graduate courses in Pattern Recognition (ECE 544) and Machine Learning (CS 446/ECE 449). Previously taught at University of Toronto and ETH Zurich. Labs & Collaborations: Leads research at Coordinated Science Laboratory (UIUC) with collaborations across University of Toronto, ETH Zurich, and industry partners like Samsung SAIT and Amazon.
Prof. Matthias Nießner is a Professor at the Technical University of Munich, leading the Visual Computing Lab. His research intersects computer graphics, vision, and AI, focusing on 3D reconstruction, semantic understanding, and AI-driven video synthesis. He holds a PhD from the University of Erlangen-Nuremberg (2013) and was a Visiting Assistant Professor at Stanford University (2013–2017). Notable awards include the ERC Starting Grant (2018), Nvidia Professorship Award, and Eurographics Young Researcher Award (2019). His work has been featured in mainstream media and led to startups like Synthesia Inc. Research spans Gaussian splatting, neural radiance fields, and generative AI for 3D avatars. Over 150 publications include SIGGRAPH, CVPR, and ECCV, with best paper awards. Projects like Face2Face and ScanNet have driven innovation in facial reenactment and 3D scene datasets. Education: PhD in Computer Science, University of Erlangen-Nuremberg (2013) Diploma in Computer Science, University of Erlangen-Nuremberg (2010) Research Interests: 3D digitization, neural rendering, generative AI, non-rigid reconstruction, and applications in AR/VR. Awards: ERC Starting Grant (2018) Nvidia Professorship Award (2018) Google Faculty Award (2018) SIGGRAPH Best Emerging Tech Award (2016) Grants: Over €1.5M from ERC and industry partnerships. Labs/Teams: Visual Computing Lab at TUM and Synthesia Inc. (co-founder). Key projects include ScanNet (large 3D indoor dataset), Face2Face (real-time facial reenactment), and Gaussian-based 3D avatars. Current work focuses on diffusion models, neural radiance fields, and AI-generated media detection.
Derek W Hoiem is a Professor in the Siebel School for Computing and Data Science at the University of Illinois Urbana-Champaign, where he has been a faculty member since 2009. His research focuses on computer vision and related areas, and he is also the co-founder and Chief Science Officer of Reconstruct, an AI-based construction technology company. His educational background includes: PhD in Robotics, Carnegie Mellon University (2007) Beckman Postdoctoral Fellowship (2008) Prof. Hoiem's research spans computer vision, with a focus on object recognition, scene understanding, and graphics. His work also extends to mobile robotics and 3D scene reconstruction. He has made significant contributions in areas such as visual recognition, 3D modeling, and the application of computer vision in construction monitoring. His recent publications (2023-2025) demonstrate a strong focus on advancing multimodal understanding, particularly in region-based representations, 3D vision, and neural radiance fields. There is a clear trend towards integrating language and vision, improving efficiency in neural networks, and applying computer vision to real-world problems such as construction progress monitoring. His scientific awards and honors are extensive and include: IEEE Fellow (2022) University Scholar (2022) Koendrink Prize (2022) Dean's Award for Excellence in Research, Associate Professor (2021) Campus Distinguished Promotion Award (2015) Best Paper Award: IEEE Winter Conference on Applications in Computer Vision (WACV) (2015) CW Gear Junior Faculty Award (2014) IEEE PAMI Young Researcher Award (2014) Dean's Award for Excellence in Research, Assistant Professor (2014) Sloan Research Fellowship (2013) Intel Early Career Faculty Honor Program Award (2012) NSF CAREER Award (2011) ACM Doctoral Dissertation Award, Honorable Mention (2008) Carnegie Mellon University SCS Distinguished Dissertation Award (2008) Best Paper Award: IEEE Computer Vision and Pattern Recognition (CVPR) (2006) Prof. Hoiem has secured significant research funding, including an NSF CAREER award and an Intel Early Career Faculty award. He is also actively involved in technology transfer, having co-founded Reconstruct where he serves as Chief Science Officer. His teaching excellence is reflected in multiple "List of Teachers Ranked as Excellent" awards spanning from 2010 to 2021. Prof. Hoiem leads a research group at UIUC focused on computer vision and 3D scene understanding. Additionally, he co-founded and serves as Chief Science Officer at Reconstruct, which develops AI-based solutions for construction monitoring.
David A. Smith is an Associate Professor at the Khoury College of Computer Sciences, Northeastern University. His research focuses on Natural Language Processing (NLP) and computational linguistics, with applications in machine translation, information retrieval, digital humanities, and social sciences. He is a founding member of the NULab for Texts, Maps, and Networks, a research center focused on digital humanities and computational social sciences. Smith's work has been funded by grants from the Mellon Foundation, NEH, and IMLS, supporting projects such as the Viral Texts initiative analyzing 19th-century newspaper networks and the Oceanic Exchanges project tracking transnational information flows. He has contributed to advancements in OCR for historical texts, text reuse detection, and computational analysis of classical languages. He has advised numerous PhD students, including Shijia Liu, Si Wu, and Ryan Muther, and teaches courses like Natural Language Processing and Information Retrieval. His research has been featured in outlets like Wired and the Economist .