Andrew Zisserman is a Royal Society Research Professor at the University of Oxford's Department of Engineering Science, affiliated with the Visual Geometry Group (VGG). His research focuses on computer vision, artificial intelligence, and neural networks, with significant contributions to multimodal learning, video understanding, and 3D scene analysis. He leads projects exploring visual-language models, audio-visual synchronization, and clinical imaging applications. Key research areas include: Video analysis and temporal modeling Multimodal systems for sign language translation and action recognition 3D shape estimation and physical property inference Foundation models and cross-modal retrieval Recent work highlights: Developed Flamingo and Tapir models for video-language tasks Advancements in spinal MRI analysis and clinical imaging Leadership in EGO4D and VoxCeleb challenges Honors include Fellowship of the Royal Society (FRS) and the ISSLS Prize in Clinical Science 2023 for spinal analysis innovations. His lab collaborates globally, emphasizing real-world applications in healthcare and autonomous systems.
Jiatao Gu is an Assistant Professor in the Department of Computer and Information Science (CIS) at the University of Pennsylvania, with a part-time role as Staff Research Scientist at Apple (MLR). He holds a Ph.D. in Electrical and Electronic Engineering from the University of Hong Kong (2018) and a B.Eng. in Electronic Engineering from Tsinghua University (2014). His research focuses on generative machine learning and AI agent interaction with the physical world, emphasizing multi-modal systems spanning language, images, videos, and 3D. Key themes include efficient modeling , flexible architecture design , and scalable decision-making frameworks . 2025: ICLR paper on DART framework 2024: TMLR work on GFlowNet alignment 2023: NeurIPS research on diffusion stability 2022: ACL papers on speech translation Recent publications explore diffusion models for text-to-image synthesis, 3D reconstruction, and efficient sampling techniques. His work addresses fundamental challenges in attention mechanisms, entropy collapse, and multi-stage distillation while advancing non-autoregressive translation and vision-language reasoning . Prospective students can apply through his recruitment process at UPenn. Prior affiliations include Meta AI (FAIR Labs) and academic collaborations with institutions like New York University's CILVR Lab.
Aishwarya Agrawal is an Assistant Professor at Université de Montréal in the Department of Computer Science and Operations Research (DIRO), affiliated with Mila – Quebec Institute of Artificial Intelligence and a Canada CIFAR AI Chair. She also serves as a research scientist at Google DeepMind, spending one day weekly there. Education: B.E. in Electrical Engineering (IIT Gandhinagar, 2014), Ph.D. in Computer Science (Georgia Tech, 2019). Her research focuses on multimodal learning , deep learning , natural language processing , and computer vision , particularly in developing AI systems that 'see' and 'communicate' effectively. Grants & Awards: Canada CIFAR AI Chair, 2020 Sigma Xi Best PhD Thesis Award, NVIDIA Fellowship (2018–2019), and multiple fellowships from Google and Facebook. She leads projects like Advancing Multimodal Vision-Language Learning (CRSNG-funded) and StarDoc: Document Structure Extraction (MITACS). Research Contributions: Pioneered benchmarks like CulturalVQA and UI-Vision , and frameworks such as PROGRESS for efficient VLM training. Her work emphasizes cross-modal alignment, robust evaluation, and cultural understanding in AI systems. Labs/Teams: Active in Mila’s core academic group and collaborates with Google DeepMind on multimodal and vision-language research. Supervises a dynamic team of PhD and master’s students in Montreal.
Subhransu Maji is an Associate Professor in the Manning College of Information and Computer Sciences at the University of Massachusetts Amherst, and the co-director of the Computer Vision Lab. He is also affiliated with the Center for Data Science and holds a part-time role as an Amazon Scholar. His research focuses on high-level visual recognition algorithms and interdisciplinary applications in ecology and astronomy. He has received prestigious awards including the NSF CAREER Award (2018), Best Paper at WACV 2015, and the Google Graduate Fellowship (2008). Education: PhD in Computer Science from UC Berkeley (2011), BTech from IIT Kanpur (2006). Prior roles include Research Assistant Professor at Toyota Technological Institute at Chicago (2012-2014). Research Interests: Computer Vision Machine Learning AI Applications in Ecology and Astronomy 3D Shape Understanding Climate Science Grants and Funding: Supported by NSF, NASA, Climate Change AI, and industry grants from Facebook, NVIDIA, Adobe, and Dolby. Current projects include satellite imagery analysis for ecology and material science applications using deep learning. Labs and Teams: Leads the Computer Vision Lab, collaborates with interdisciplinary teams on ecological monitoring (e.g., bird migration tracking via radar data) and material property prediction (e.g., zeolite adsorption modeling).
Deva Ramanan is a Professor at the Robotics Institute of Carnegie Melllon University, where he leads research in computer vision and machine learning. His work focuses on modeling human visual perception, leveraging large-scale visual data, and developing systems for 3D understanding, neural rendering, and autonomous systems. He advises a large group of PhD students and has mentored numerous postdoctoral researchers now in leading roles across industry and academia. His research interests include computer vision, machine learning, human perception modeling, 3D scene understanding, neural rendering, autonomous driving, video understanding, and multimodal foundation models. These areas reflect his focus on both foundational models and their application to real-world problems in robotics and AI. The recent publications highlight a strong trend toward multimodal and 3D-aware models, with increasing use of diffusion models, neural fields, and large vision-language systems. Key themes include scene flow, 3D reconstruction from monocular video, autonomous driving perception, and robust evaluation of vision-language models. There is a clear emphasis on both methodological innovation and practical deployment in dynamic environments. Marr Prize, Honorable Mention (ICCV 2021) Best Paper, Honorable Mention (ECCV 2020) Best Paper Finalist (WACV 2024) Best Paper Award (WACV 2016) Best Industrial Paper, Honorable Mention (BMVC 2017) Marr Prize winner (ICCV 2009) Deva Ramanan has advised numerous PhD and master’s students, many of whom are now at top institutions and companies including Apple, Meta, Google, Nvidia, OpenAI, and Princeton. He has received substantial funding from IARPA, DARPA, NSF, Intel, Google, and Facebook for projects in video analytics, dispersed computing, visual cloud systems, and multi-task recognition. His group has developed influential datasets and benchmarks used widely in the community. He leads a vibrant research lab focused on advancing computer vision through deep learning and multimodal integration. His team works on core challenges in perception, including 3D reconstruction, motion modeling, object detection, and scene understanding, with applications in robotics and autonomous systems.
Andreas Vlachos is a Professor of Natural Language Processing and Machine Learning at the Department of Computer Science and Technology, University of Cambridge, and holds the Dinesh Dhamija Fellowship at Fitzwilliam College. His research spans dialogue modeling, automated fact-checking, imitation learning, semantic parsing, biomedical text mining, and trustworthiness in AI systems. PhD in Computer Science, University of Cambridge (supervised by Ted Briscoe and Zoubin Ghahramani) Lecturer at University of Sheffield Postdoctoral roles at UCL, University of Cambridge (NLIP group, Stephen Clark), and University of Wisconsin-Madison (Mark Craven) Current research focuses on evaluating and mitigating biases in language models, advancing fact-checking methodologies, and improving model robustness through interpolation, reinforcement learning, and causal reasoning. His work integrates natural logic, knowledge graphs, and multimodal evidence for verification tasks. Recent publications address uncertainty quantification, temporal planning benchmarks, and ethical framing of NLP artifacts. Grants from ERC, EPSRC, Facebook, Google, and the Alan Turing Institute fund his research team. Collaborations include Sebastian Riedel, Stephen Clark, and Mark Craven. Key projects explore disinformation detection, long-form generation, and confidence calibration in AI systems.
Almut Sophia Koepke is a junior research group leader at the Technical University of Munich and University of Tübingen, focusing on multimodal learning problems integrating sound, vision, and text. Her work bridges foundational research in audio-visual understanding with practical applications in few-shot learning, zero-shot translation, and cross-modal attention mechanisms.
Zaid Harchaoui is an Adjunct Professor in the Department of Statistics at the University of Washington. His research focuses on machine learning, generative AI, and algorithmic optimization, with applications spanning ecology, neuroscience, and artificial intelligence. He explores learning under distributional shifts and develops tools for scalable generative models in language and vision domains. University: University of Washington Department: Statistics Research Focus: Learning from data with computational, inferential, and mathematical rigor; distributional shift adaptation; generative model scaling Email: zaid@uw.edu His recent work emphasizes generative AI applications in ecology and neuroscience, stochastic optimization for robustness, and algorithmic efficiency in large-scale learning. Key contributions include techniques for distributionally robust optimization, interpretable authorship obfuscation, and uncertainty quantification in behavior classification. Scientific awards and honors are not explicitly mentioned in the provided text. Collaborative efforts often intersect with nonlinear control algorithms, spectral analysis, and privacy-preserving machine learning frameworks.
Rana Hanocka is an Assistant Professor of Computer Science at the University of Chicago, leading the 3DL research group focused on AI-driven 3D geometry processing. She earned her Ph.D. in 2021 from Tel Aviv University under Professors Daniel Cohen-Or and Raja Giryes. Her work explores neural networks for unstructured 3D data, including mesh convolutional networks and self-priors for shape reconstruction. Key research areas include geometric deep learning, human-AI collaboration in 3D modeling, and interpretable neural networks. Notable contributions include MeshCNN (SIGGRAPH 2019) and Point2Mesh (NeurIPS 2020), advancing mesh analysis and point cloud processing. Awards include the 2023 Pazy Research Award and 2020 Rising Star in EECS. Education : Ph.D. Computer Science, Tel Aviv University (2021) Labs/Groups : 3DL Group at UChicago Grants : NSF Grant for AI-driven 3D modeling tools (2023) Research directions emphasize creative human-AI partnerships, unsupervised learning from shape collections, and explainable 3D neural networks. Ongoing projects include style-aware 3D synthesis, interactive segmentation, and multimodal shape interfaces.
Oliver Kroemer is an Associate Professor at Carnegie Mellon University's Robotics Institute (RI), affiliated with the Intelligent Autonomous Manipulation (IAM) Lab. His research focuses on enabling robots to learn versatile manipulation skills through lifelong frameworks, with applications in elder care, environmental maintenance, and hazardous operations. Developed methods for robot learning via physical interaction and reinforcement learning Created representations for contact states and motor primitives to improve skill generalization Research Interests: Spanning robot learning, tactile sensing, force-velocity control, and lifelong skill acquisition. Projects include Agile and Dynamic Interactions for Mobile Manipulation and Integrated Planning and Learning (Pillar project). Scientific Awards: Finalist, Georges Giralt Ph.D. Award (2015) Education: Masters & Bachelors in Engineering, University of Cambridge (2008) Ph.D., Technische Universitaet Darmstadt (2014) Students & Affiliates: Current PhD: Mark Lee, Sarvesh Patil, Saumya Saxena, Yunus Seker, Zilin Si Past PhD: Alex LaGrassa, Tabitha Lee, Qiao Liang, Shivam Vats, Kevin Zhang
Matthias Hein is a Professor at the Department of Computer Science, Faculty of Mathematics and Natural Sciences, University of Tübingen. His research focuses on Machine Learning , Adversarial Robustness , and Out-of-Distribution Detection , with applications in computer vision and medical imaging. He has received notable recognition including the Best Paper Honorable Mention Prize at ICLR 2021 and Outstanding Paper Award at CVPR 2021. His work includes developing benchmarks like RobustBench and Spurious ImageNet , and frameworks such as Sparse-RS and DIG-IN . His recent publications emphasize adversarial robustness across multiple domains (vision, text), counterfactual explanations for classifiers, and improved OOD detection methods . Collaborators include prominent researchers like Francesco Croce, Julian Bitterwolf, and Alexander Meinke. Scientific Awards : Best Paper Honorable Mention (ICLR 2021) CVPR 2021 Outstanding Paper Award Key Research Areas : Adversarial Robustness Vision-Language Models Medical Imaging AI Neural Network Calibration
Dai Zhongxiang is an Assistant Professor and Presidential Young Fellow at the School of Data Science, The Chinese University of Hong Kong, Shenzhen (CUHKSZ), where he joined in August 2024. Previously, he was a Postdoctoral Associate at MIT's Laboratory for Information and Decision Systems (January-June 2024) and a Postdoctoral Fellow at the National University of Singapore's Department of Computer Science (April 2021-December 2023). He completed his Ph.D. in Artificial Intelligence at NUS under the supervision of Bryan Kian Hsiang Low and Patrick Jaillet. Dr. Dai's research focuses on the intersection of theoretical and practical AI, with particular emphasis on large language models (LLMs) and optimization techniques. His work spans both theoretical foundations of multi-armed bandits and Bayesian optimization, as well as practical applications in LLM inference, including prompt optimization, in-context learning, personalization of LLMs, LLM-based agents, and scaling up test-time computation of LLMs. His research approach often bridges theoretical principles with real-world applications, particularly in AI4Science problems. His recent publications demonstrate a clear trend toward advancing LLM capabilities through optimization techniques, with increasing focus on practical deployment challenges. The research spans both theoretical contributions to optimization theory and applied work on enhancing LLM performance in real-world scenarios. His work on dueling bandits, neural bandits, and zeroth-order optimization has been consistently published in top-tier venues including NeurIPS, ICML, ICLR, and ACL. Presidential Young Fellow, CUHKSZ (2024) Dean's Graduate Research Excellence Award, NUS (2021) Research Achievement Award × 2, NUS (2019 & 2020) Singapore-MIT Alliance Graduate Fellowship (2017) Dr. Dai actively mentors multiple Ph.D. students and research assistants, with several of his students' papers accepted to top conferences. His research has received significant attention, with invitations to serve as Area Chair for NeurIPS 2025 and ICLR 2025, reflecting his growing influence in the machine learning community. His work bridges theoretical machine learning with practical applications in large-scale AI systems.
WANG Ye is an Associate Professor in the Department of Computer Science at the School of Computing, National University of Singapore (NUS). He holds a PhD in Information Technology from Tampere University of Technology, Finland, and has been a tenured faculty member at NUS since 2002, following his industry research role at Nokia Research Center. He is the director of the Sound and Music Computing Lab at NUS, leading cutting-edge research in AI-driven music and health technologies. PhD, Information Technology, Tampere University of Technology, Finland (2002) MSc, Telecommunications, Braunschweig University of Technology, Germany (1993) BSc, Telecommunications, South China University of Technology, China (1983) His research is centered on Sound and Music Computing for Human Health and Potential (SMC4HHP) , with a focus on eHealth, eLearning, mobile/wearable computing, and music information retrieval. His work spans AI for stroke rehabilitation, language learning through singing, singing voice synthesis, and automatic music transcription. He has pioneered systems like SLIONS (language learning via karaoke), CocoLyricist (AI co-creation for stroke recovery), and SinTechSVS (expressive singing voice synthesis). The latest articles highlight a strong trend in AI-driven music and health technologies , particularly in controllable lyric generation, singing voice synthesis, automatic pronunciation assessment, and multimodal music transcription. The research increasingly integrates large language models, explainable AI, fairness, and real-world deployment, reflecting a shift from theoretical exploration to practical, human-centered applications in healthcare and education. Dr. Wang has received numerous scientific honors, including: Best Paper Awards at ACM MM, ISMIR, IEEE ISM, and CHI First Prize, Asia Pacific Assistive, Rehabilitative, and Therapeutic Technologies Challenge (2015) Faculty Teaching Excellence Award, NUS School of Computing (2024) Top Paper Award, ACM Multimedia 2022 AI in Medicine Collaborative Grant for CocoLyricist project He has supervised over 11 PhD and 20 MComp students and is currently guiding six PhD candidates. His grants come from MOE, NRF, A*STAR, Nokia, and Smule. He has served as General Chair of ISMIR2017 and TPC Co-Chair of ICOT2017, and is on the editorial boards of IEEE Transactions on Multimedia and Journal of New Music Research. He has also developed and taught the first course on Sound and Music Computing in Singapore. Dr. Wang leads the Sound and Music Computing Lab (SMC Lab) , a multidisciplinary team exploring the synergy of music computing, AI, mobile technology, and cloud systems for health and education. The lab actively collaborates with medical institutions such as NUS Yong Loo Lin School of Medicine, Singapore General Hospital, and Harvard Medical School, and is currently working on projects in AI-supported language learning, stroke rehabilitation, and intelligent music interfaces.
Shubham Tulsiani is an Assistant Professor at Carnegie Mellon University's Robotics Institute, where he leads the Computer Vision group and the Physical Perception Lab. His research focuses on inferring physically and spatially grounded representations from perceptual inputs, with applications in 3D vision, robot manipulation, and neural scene reconstruction. He directs an active research group with multiple PhD and Master's students. Research interests center on 3D scene understanding , robot learning , and generative modeling , with specific emphasis on: self-supervised perception, neural rendering, multi-view geometry, manipulation from visual inputs, and physics-based reasoning. The lab develops methods that leverage physical world constraints as supervisory signals. Recent publications demonstrate strong focus on diffusion models for 3D tasks , sparse-view reconstruction , and robotic manipulation transfer . Key trends include neural inverse rendering, view synthesis from limited observations, and translating human interactions to robot actions. Awards include: Best Student Paper Award at CVPR 2015 Advising includes supervision of 5 PhD students, 4 MS students, and undergraduates. Lab alumni hold positions at Google, Stanford, Meta, and Princeton. The Physical Perception Lab collaborates with FAIR Pittsburgh and the CMU Computer Vision group.
Timothy M. Hospedales is a Professor of Artificial Intelligence at the Institute of Perception, Action and Behaviour within the School of Informatics at the University of Edinburgh . He also serves as VP AI and Head of Samsung AI Research Centre Europe . His research focuses on efficient and robust AI , emphasizing meta-learning , lifelong transfer-learning , and domain adaptation in both probabilistic and deep learning frameworks. Applications span computer vision , vision and language , reinforcement learning for robotics , and finance . Professor at University of Edinburgh (2020–present) ELLIS Fellow (2021) Head of Samsung AI Research Europe (2020–present) Founding Director of Applied Machine Learning Lab at QMUL (2012–2016) His work includes pioneering contributions to meta-learning , few-shot learning , and self-supervised methods , with notable awards such as the Best Paper Prize at ICML AutoML 2018 and Best Student Paper at ICPR 2018 . He has co-authored 15+ recent papers on topics like Vision-Language Models , Medical AI Fairness , and Diffusion Model Optimization . He served as Program Co-Chair for BMVC 2018 and AAAI 2022 , and authored a book on Visual Adaptation in the Deep Learning Era (2022). Co-Chair, BMVC 2018 Guest Editor, IET CV Special Issue (2016) Keynote Speaker at TASK-CV Workshop (ECCV 2016) Special Issue on Fewer Labels (IEEE PAMI 2020) His leadership extends to organizing workshops like the Learning-to-Learn Workshop at ICLR 2021 , Meta-Learning Workshop at NeurIPS 2020 , and Domain Generalisation Workshop at ICLR 2023 . Current projects include Meta-Omnium (CVPR 2023) for general-purpose meta-learning and MetaAudio (ICANN 2022) for few-shot audio classification benchmarks.