Ajmal Mian is a Professor of Computer Science at the University of Western Australia (UWA), affiliated with the School of Physics, Maths and Computing. He holds an Australian Research Council Future Fellowship (2022) and leads research in Artificial Intelligence, Computer Vision, and Machine Learning. His work focuses on 3D computer vision, adversarial AI defense, and explainable AI. His research interests include 3D point cloud analysis, face recognition, human action recognition, and remote sensing. He has published over 300 papers and secured major grants from ARC, NHMRC, and DARPA, totaling millions in funding. He has supervised 29 PhD students and mentored 12 postdoctoral researchers. Key projects include 3D diffusion models for scene generation, robust 3D vision systems, and defense against AI deception attacks. He serves as a fellow of IAPR, an ACM Distinguished Speaker, and has editorial roles at IEEE Transactions on Neural Networks and Pattern Recognition. Research Awards: HBF Mid-Career Scientist of the Year, West Australian Early Career Scientist of the Year, IAPR Best Scientific Paper Award. Grants: ARC Discovery Projects, National Intelligence & Security Discovery grants, DARPA grants for AI security. His teaching spans computer vision, machine learning, and programming courses. Collaborations include defense, medical, and agricultural applications.
Alexei A. Efros is the Howard Friesen Professor in the EECS Department at the University of California, Berkeley, and a core member of the Berkeley Artificial Intelligence Research (BAIR) Lab. Previously, he spent a decade at Carnegie Mellon University’s Robotics Institute. His research focuses on data-driven computer vision, self-supervised learning, computational photography, and generative models. He has pioneered advancements in visual representation learning, including seminal work on neural radiance fields and generative adversarial networks. Education Background: Efros holds a PhD in Computer Science from MIT, though specific details of his academic journey are not explicitly provided in the text. His career includes postdoctoral research at the University of Oxford with Andrew Zisserman and collaborative work with Team WILLOW at INRIA Paris. Research Interests: Efros explores how vast uncurated visual data can be leveraged for understanding and synthesizing the visual world. Key areas include self-supervised learning, generative models, and applications in robotics and art. His lab has contributed influential techniques such as Style Transfer, GAN-based image synthesis, and neural scene representation learning. Recent work emphasizes real-time adaptation (Test-Time Training), 3D perception models, and ethical AI implications of generative systems. Publications: Over 150+ publications span topics like Generative Adversarial Networks (GANs), unsupervised learning, and visual-linguistic models. Notable works include Unpaired Image-to-Image Translation (CUT/GAU), Style Transfer , and Swapping Autoencoder . His research has significant industry impact, with techniques adopted in Adobe’s software and generative AI applications. Grants & Collaborations: Efros has secured major funding from NSF, DARPA, and industry partnerships (e.g., Adobe, NVIDIA). He co-leads projects on scalable vision models, ethical AI, and real-world perception systems. Current collaborations include work with MIT, NYU, and INRIA Paris. Labs & Teams: Leads the BAIR Vision Group at Berkeley, fostering interdisciplinary research between computer vision, graphics, and robotics. The group emphasizes Slow Science principles, prioritizing deep exploration over rapid publication.
Huaxiu Yao is an Assistant Professor at the University of North Carolina at Chapel Hill, holding a joint appointment in the School of Data Science and Society and the Department of Computer Science (College of Arts & Sciences). His research focuses on building reliable large-scale AI models (foundation models) with applications in healthcare, robotics, genomics, and transportation. He leads the AIMING Lab, which explores adaptive intelligence through alignment, interaction, and learning. Education: Ph.D. from Pennsylvania State University (2021), Postdoctoral Scholar at Stanford University (hosted by Chelsea Finn). Research Interests: Generalizable AI agents, preference alignment, out-of-distribution generalization, embodied AI, and multimodal reasoning. Key applications include biomedicine, robotics, and vision-language systems. Notable Awards: KDD Best Paper Award (2024), Amazon Research Awards (2025), TMLR Outstanding Paper Award (2024). Advising & Labs: Recruits Ph.D. and intern students. Leads the AIMING Lab, affiliated with UNC NLP Group. Organizes workshops on foundation models (ICML 2024) and trustworthy AI systems. Publications: Over 40 peer-reviewed papers, including top venues like ICLR, NeurIPS, and ICML. Focuses on AI alignment, multimodal systems, and domain generalization.
Matthew O'Toole is an Associate Professor at Carnegie Mellon University's School of Computer Science, holding joint appointments in the Robotics Institute and Computer Science Department. His research focuses on computational imaging, integrating optics, electronics, and computational processing to innovate visual information capture and display. Education: PhD (Computer Science, University of Toronto, 2016), MSc (2009), BSc (Honors Computer Science and Mathematics, University of British Columbia, 2007). Prior roles include Banting Postdoctoral Fellow at Stanford University and visiting scholar at MIT Media Lab's Camera Culture group. Research interests emphasize programmable imaging systems, transient imaging, non-line-of-sight sensing, and holographic displays. Key innovations include vibration sensing via dual-shutter optics and radar super-resolution for autonomous vehicles. Awards include runner-up best paper recognitions at ICCV 2007, CVPR 2014, and SIGGRAPH 2017 dissertation honors. Advisees include Dorian Chan and Arjun Teh. Grants supported by Canadian Banting Fellowships. Active in workshop organization (CVPR Computational Cameras 2016-2017) and course development on computational imaging at SIGGRAPH 2014. Labs/Teams: Leads research in computational imaging and robotics at CMU, collaborating with industry partners like NVIDIA and MDA. Current projects explore LiDAR-radar fusion, holographic projection systems, and dynamic scene reconstruction.
Shubham Tulsiani is an Assistant Professor at Carnegie Mellon University's Robotics Institute, where he leads the Computer Vision group and the Physical Perception Lab. His research focuses on inferring physically and spatially grounded representations from perceptual inputs, with applications in 3D vision, robot manipulation, and neural scene reconstruction. He directs an active research group with multiple PhD and Master's students. Research interests center on 3D scene understanding , robot learning , and generative modeling , with specific emphasis on: self-supervised perception, neural rendering, multi-view geometry, manipulation from visual inputs, and physics-based reasoning. The lab develops methods that leverage physical world constraints as supervisory signals. Recent publications demonstrate strong focus on diffusion models for 3D tasks , sparse-view reconstruction , and robotic manipulation transfer . Key trends include neural inverse rendering, view synthesis from limited observations, and translating human interactions to robot actions. Awards include: Best Student Paper Award at CVPR 2015 Advising includes supervision of 5 PhD students, 4 MS students, and undergraduates. Lab alumni hold positions at Google, Stanford, Meta, and Princeton. The Physical Perception Lab collaborates with FAIR Pittsburgh and the CMU Computer Vision group.
Timothy M. Hospedales is a Professor of Artificial Intelligence at the Institute of Perception, Action and Behaviour within the School of Informatics at the University of Edinburgh . He also serves as VP AI and Head of Samsung AI Research Centre Europe . His research focuses on efficient and robust AI , emphasizing meta-learning , lifelong transfer-learning , and domain adaptation in both probabilistic and deep learning frameworks. Applications span computer vision , vision and language , reinforcement learning for robotics , and finance . Professor at University of Edinburgh (2020–present) ELLIS Fellow (2021) Head of Samsung AI Research Europe (2020–present) Founding Director of Applied Machine Learning Lab at QMUL (2012–2016) His work includes pioneering contributions to meta-learning , few-shot learning , and self-supervised methods , with notable awards such as the Best Paper Prize at ICML AutoML 2018 and Best Student Paper at ICPR 2018 . He has co-authored 15+ recent papers on topics like Vision-Language Models , Medical AI Fairness , and Diffusion Model Optimization . He served as Program Co-Chair for BMVC 2018 and AAAI 2022 , and authored a book on Visual Adaptation in the Deep Learning Era (2022). Co-Chair, BMVC 2018 Guest Editor, IET CV Special Issue (2016) Keynote Speaker at TASK-CV Workshop (ECCV 2016) Special Issue on Fewer Labels (IEEE PAMI 2020) His leadership extends to organizing workshops like the Learning-to-Learn Workshop at ICLR 2021 , Meta-Learning Workshop at NeurIPS 2020 , and Domain Generalisation Workshop at ICLR 2023 . Current projects include Meta-Omnium (CVPR 2023) for general-purpose meta-learning and MetaAudio (ICANN 2022) for few-shot audio classification benchmarks.
Karthik R. Narasimhan is a Professor at Princeton University's School of Engineering and Applied Science in the Department of Computer Science. Previously, he earned his PhD from MIT under Regina Barzilay and served as a visiting research scientist at OpenAI during 2017-18. His research focuses on the intersection of language and decision-making, building autonomous agents that learn from both experience and human knowledge. His research spans multiple high-impact areas including language agents (Text-DQN, CALM, ReAct, Tree of Thoughts), reinforcement learning (h-DQN, Multi-Objective RL), and AI safety (Toxicity in ChatGPT, DataMUX). He has developed critical datasets and benchmarks such as WebShop, InterCode, SWE-bench, and SILG that have become standard evaluation tools in the field. Current work emphasizes agent capabilities, software engineering automation, and multimodal interaction. His publication trends show strong focus on practical agent deployment (SWE-agent, Tree of Thoughts), safety evaluation (Probing AI Safety), and efficiency improvements (DataMUX). Recent work increasingly addresses real-world challenges in software engineering, security, and human-AI collaboration through rigorous benchmarking. Co-author of foundational GPT (2018) paper Key developer of Text-DQN (2015), CALM (2020), ReAct (2022), Tree of Thoughts (2023) Creator of influential benchmarks: WebShop (2022), SWE-bench (2023), InterCode (2023) He actively advises students through Princeton's computer science program, with research supported by multiple grants focused on autonomous agent development and language-based decision systems. His GitHub repositories (nlp-datasets, text-world-player) demonstrate strong community engagement in open-source research tools. Current projects include advancing language agent capabilities through Reflexion (2023) and Tree of Thoughts (2023) frameworks while addressing critical safety and efficiency challenges.
Prof. Matthias Nießner is a Professor at the Technical University of Munich , where he leads the Visual Computing Lab . Prior to this, he held a Visiting Assistant Professor position at Stanford University . His work bridges computer vision , graphics , and machine learning , focusing on 3D reconstruction , semantic scene understanding , and AI-driven video synthesis . Prof. Nießner has published over 150 works in top venues like SIGGRAPH , CVPR , and ECCV , with several receiving best paper awards (SIGCHI’14, HPG’15, SPG’18, SIGGRAPH’16 Emerging Tech). His research has garnered international media attention, including features in the New York Times , Wall Street Journal , and MIT Technological Review , as well as TV demonstrations (e.g., Jimmy Kimmel Live for Face2Face technology). Awards : TUM-IAS Rudolph Moessbauer Fellowship (2017–ongoing) Google Faculty Award (2017) Nvidia Professor Partnership Award (2018) ERC Starting Grant (2018, €1.5M) Eurographics Young Researcher Award (2019) Research Trends : 3D Gaussian Splatting for real-time rendering Neural Radiance Fields (NeRF) with mesh supervision Audio-driven facial animation via diffusion models Latent space diffusion for 3D scenes Self-supervised and zero-shot methods for 3D and image analysis As a co-founder and director of Synthesia Inc. , he drives democratization of synthetic media. His YouTube channel has over 5 million views, reflecting his impact beyond academia.
Animesh Garg is an Assistant Professor at the School of Interactive Computing at Georgia Tech, where he leads the People, AI, and Robotics (PAIR) research group . He holds a Senior Researcher position at Nvidia Research and has courtesy appointments at the University of Toronto and Vector Institute. Previously, he served as Chief Scientific Officer at Apptronik (2024-2025) and Senior Staff Research Scientist at Nvidia Research (2018-2024). Education : Ph.D. in Operations Research from UC Berkeley (2011-2016), MS in Computer Science and Industrial Engineering from Georgia Tech and University of Delhi. Research Focus : Building Generalizable Autonomy through Reinforcement Learning , Control Theory , and 3D Vision , with applications in Surgical Robotics , Self-Driving Labs , and Manufacturing . Key Article Themes : His recent work emphasizes Foundation Models for robotics, Differentiable Simulation , Language-Guided Autonomy , and Structured Inductive Biases in sequential decision-making. Scientific Awards : Stephen Fleming Early Career Professorship at Georgia Tech. Teaching : Courses on AI, Deep Reinforcement Learning, and Algorithmic Intelligence in Robotics at Georgia Tech. Labs & Collaborations : Affiliated with Institute for Robotics and Intelligent Machines (IRIM) and ML@GT at Georgia Tech; collaborates intensively with Nvidia Robotics.
Andreas Geiger is a Professor and Head of the Department of Computer Science at the University of Tübingen, Germany. He leads the Autonomous Vision Group (AVG) within CyberValley and is a core faculty member of the Tübingen AI Center. His roles also include PI in the ML in Science Excellence Cluster and the CRC Robust Vision, as well as ELLIS Fellow and coordinator of the ELLIS PhD program. He specializes in machine learning models for computer vision, robotics, and autonomous systems, with applications in self-driving cars, VR/AR, and scientific document analysis. Educational background: While not explicitly detailed, his positions imply a Ph.D. in Computer Science or related field. His work spans interdisciplinary collaborations with institutions like ETH Zürich, Microsoft, and the University of Bonn. Research focuses on 3D scene understanding, Gaussian splatting, generative models, and reliable autonomous systems. Notable contributions include the KITTI dataset and foundational work in neural radiance fields. Awards include the Sage 10-Year Impact Award (2024), ERC Starting Grant (2019), and IEEE PAMI Young Researcher Award (2018). Key projects include the Scholar Inbox paper recommender platform, ReSim (reliable world simulation), and advancements in 3D scene generation (e.g., UrbanCAD, PrITTI). His lab maintains a strong focus on open-source tools and datasets, such as the CARLA Route Generator. Grants and funding include support from Vector Stiftung (MINT innovation program) and EU initiatives like the ML in Science Cluster. His team collaborates internationally, with recent work presented at CVPR, SIGGRAPH, and NeurIPS.
Yonatan Bisk is an Assistant Professor at Carnegie Mellon University within the Language Technologies Institute (with courtesy appointment in Robotics Institute). His research bridges Natural Language Processing , Robotics , and Embodied AI , focusing on language grounding, theory of mind, and multimodal interaction. Education : Ph.D. in Computer Science from University of Illinois at Urbana-Champaign Postdoctoral Experience : USC ISI, University of Washington, Allen Institute for AI Industry Appointments : Microsoft Research, Meta AI His research emphasizes embodied language systems and social intelligence in AI . Recent projects include WebArena for autonomous agents, SOTOPIA for social reasoning, and HomeRobot for open-vocabulary manipulation. He leads the REAL Center (Robotics, Embodied AI, and Learning) to foster interdisciplinary collaboration. Key scientific awards include selection for the DARPA ISAT Study Group (2024). He teaches courses like "Talking to Robots" and "Multimodal Machine Learning" while serving as area chair/editor across NLP, Robotics, and ML communities.
Vicente Ordóñez-Román is an Associate Professor in the Department of Computer Science at Rice University, part of the George R. Brown School of Engineering. His research focuses on the intersection of computer vision, natural language processing, and machine learning, with an emphasis on fair, transparent, and interpretable AI. He leads the Vision, Language, and Learning Lab and contributes to the Ken Kennedy Institute's Closed-loop Computer Vision research cluster. Education: PhD in Computer Science (UNC Chapel Hill, 2015), MS in Computer Science (Stony Brook University), and Engineering (Escuela Superior Politécnica del Litoral, Ecuador). Prior roles include Assistant Professor at the University of Virginia (2016-2021) and visiting positions at Adobe Research, the Allen Institute for AI, and Amazon. Research Interests : Developing multimodal AI systems that integrate visual and textual data, mitigating biases in AI, and advancing generative models. His work emphasizes ethical AI and societal impact, as seen in his contributions to the whitepaper advocating for federal regulation of facial recognition technologies. Awards & Recognition : NSF CAREER Award (2021), Marr Prize (ICCV 2013), Best Paper at EMNLP 2017, and multiple industry grants from Google, Amazon, and Facebook. His research has been featured in media outlets like WIRED, The New York Times, and Bloomberg News. Advising & Grants : Supervises a diverse research group spanning PhD, MS, and undergraduate students. Secured over $1.8 million in external funding, including NSF grants, Amazon FAI awards, and Google Cloud credits. Leads initiatives on bias mitigation, AI ethics, and multimodal learning. Labs & Collaborations : Directs the Vision, Language, and Learning Lab (vislang.ai), collaborating with industry partners like Adobe, Amazon, and SAP. Engages in interdisciplinary projects at the Ken Kennedy Institute, focusing on closed-loop computer vision systems.
Fei-Fei Li is the Sequoia Capital Professor in Computer Science at Stanford University and Founding Co-Director of the Stanford Institute for Human-Centered AI (HAI). She holds courtesy appointments in the Graduate School of Business and is a Senior Fellow at HAI. Her work bridges AI research with interdisciplinary applications in healthcare, robotics, and policy. Dr. Li pioneered the ImageNet dataset, instrumental in the AI revolution, and co-founded World Labs to advance spatial intelligence and generative AI. Education: B.A. in Physics, Princeton University (1999) Ph.D. in Electrical Engineering, Caltech (2005) Doctorate (Honorary), Harvey Mudd College (2022) Research Interests: AI ethics, computer vision, robotic learning, healthcare applications, and human-AI collaboration. Her teams developed frameworks like MOMA for activity recognition and BEHAVIOR for embodied AI benchmarks. She advocates for diversity in tech and co-founded AI4All to mentor underrepresented students. Key Contributions: ImageNet and ImageNet Challenge Stanford Vision and Learning Lab (SVL) Policy advisory roles for U.S. Senate, UN Secretary-General, and California Governor Labs & Initiatives: Leads the People, AI & Robots Group (PAIR), Partnership in AI-Assisted Care (PAC), and the Human-Centered AI Institute. Her work emphasizes ethical AI deployment and societal impact. Awards: VinFuture Prize (2024), IEEE Fellow, National Academy memberships (Engineering, Medicine, Arts & Sciences), and recognition as one of Time’s AI100 Influencers.
Andrea Tagliasacchi is an Associate Professor at Simon Fraser University's School of Computing Science, holding the Visual Computing Research Chair. He is also a part-time (20%) staff research scientist at Google DeepMind (Toronto) and an associate professor (status-only) at the University of Toronto's computer science department. His research focuses on 3D visual perception at the intersection of computer vision, graphics, and machine learning. Education: EPFL – Postdoc Simon Fraser University – PhD (NSERC Alexander Graham Bell Fellow) Politecnico di Milano – MSc (Gold Medalist) Research Interests: His work emphasizes 3D reconstruction, neural fields, and applications in robotics, autonomous systems, and augmented reality. Recent advancements include scalable 3D Gaussian splatting, robust neural rendering techniques, and diffusion models for 4D generation. Notable Articles: Recent work spans real-time differentiable ray tracing, stochastic rasterization for 3D Gaussian splats, and generative image composition using neural fields. His publications often blend theoretical contributions with practical applications in CVPR, SIGGRAPH, and NeurIPS. Awards: 2015 SGP Best Paper Award 2020 CVPR Best Student Paper Award 2024 CVPR Best Paper Honorable Mention Advising & Grants: Advised 14+ PhD/MSc students (e.g., Baptiste Angles, Sara Sabour) and co-advised with notable figures like Geoffrey Hinton. Active in grants involving neural field compression, robotic perception, and generative AI. Labs & Teams: Leads a lab at SFU focused on 3D vision and neural fields, collaborating with industry partners like Google Brain and Samsung Research.
Professor Gabriel Brostow is a faculty member in the Department of Computer Science at University College London (UCL), where he leads research in Computer Vision and Human-Computer Interaction. He also serves as Chief Research Scientist and Senior Director of the R&D Team at Niantic, the company behind Pokémon GO. His work bridges academic research and industry applications, focusing on developing AI systems that enhance human capabilities through what he terms 'Human in the Loop AI'—now commonly referred to as Human-Centered AI. Brostow completed his BS in Electrical Engineering at UT Austin, followed by a PhD with Irfan Essa at Georgia Tech. He then pursued postdoctoral research with Roberto Cipolla's Computer Vision & Robotics Group at Cambridge University as a Marshall Sherfield Fellow, and with Marc Pollefeys in ETH Zurich's CVG Group. His research explores how AI, particularly Computer Vision, can serve as 'super-tools' for professionals across various domains including filmmaking, architecture, robotics, and scientific research. Specific interests include assistive technology for everyday life, authoring systems that maximize user effort, 3D reconstruction, depth estimation, and vision-language models. His work often involves creating systems that are validated through real-world human interaction to ensure practical utility. Analysis of his recent publications reveals a strong focus on practical applications of Computer Vision that directly interact with humans. His research spans 3D scene understanding, depth estimation, sketch-based interfaces, and multimodal AI systems. There's a clear emphasis on creating benchmarks and tools that facilitate human-AI collaboration, with applications in assistive technology, urban planning, filmmaking, and biodiversity monitoring. His work frequently appears at top conferences including CVPR, NeurIPS, ECCV, and CHI. Marshall Sherfield Fellowship Brostow actively mentors PhD students, with current advisees including Ross Murphy, Skanda Koppula, Gizem Unlu, Omiros Pantazis, and Jamie Watson. His alumni include numerous PhD graduates and MSc students who have gone on to successful careers in academia and industry. He emphasizes selecting students based on passion and potential rather than just academic credentials, valuing traits like helpfulness, drive, and hunger to learn. His research is supported through collaborations with major institutions and companies including DeepMind, MIT, and the University of Edinburgh. He leads a research group at UCL that collaborates closely with Niantic's R&D team, creating a unique bridge between academic research and industry application. His team's work frequently involves developing novel Computer Vision techniques that are validated through real-world human interaction, ensuring practical utility alongside technical innovation. The group explores blue-sky research problems with applications ranging from assistive technology to professional tools for filmmakers, architects, and scientists studying diverse environments.