Andrew Zisserman is a Royal Society Research Professor at the University of Oxford's Department of Engineering Science, affiliated with the Visual Geometry Group (VGG). His research focuses on computer vision, artificial intelligence, and neural networks, with significant contributions to multimodal learning, video understanding, and 3D scene analysis. He leads projects exploring visual-language models, audio-visual synchronization, and clinical imaging applications. Key research areas include: Video analysis and temporal modeling Multimodal systems for sign language translation and action recognition 3D shape estimation and physical property inference Foundation models and cross-modal retrieval Recent work highlights: Developed Flamingo and Tapir models for video-language tasks Advancements in spinal MRI analysis and clinical imaging Leadership in EGO4D and VoxCeleb challenges Honors include Fellowship of the Royal Society (FRS) and the ISSLS Prize in Clinical Science 2023 for spinal analysis innovations. His lab collaborates globally, emphasizing real-world applications in healthcare and autonomous systems.
Almut Sophia Koepke is a junior research group leader at the Technical University of Munich and University of Tübingen, focusing on multimodal learning problems integrating sound, vision, and text. Her work bridges foundational research in audio-visual understanding with practical applications in few-shot learning, zero-shot translation, and cross-modal attention mechanisms.
WANG Ye is an Associate Professor in the Department of Computer Science at the School of Computing, National University of Singapore (NUS). He holds a PhD in Information Technology from Tampere University of Technology, Finland, and has been a tenured faculty member at NUS since 2002, following his industry research role at Nokia Research Center. He is the director of the Sound and Music Computing Lab at NUS, leading cutting-edge research in AI-driven music and health technologies. PhD, Information Technology, Tampere University of Technology, Finland (2002) MSc, Telecommunications, Braunschweig University of Technology, Germany (1993) BSc, Telecommunications, South China University of Technology, China (1983) His research is centered on Sound and Music Computing for Human Health and Potential (SMC4HHP) , with a focus on eHealth, eLearning, mobile/wearable computing, and music information retrieval. His work spans AI for stroke rehabilitation, language learning through singing, singing voice synthesis, and automatic music transcription. He has pioneered systems like SLIONS (language learning via karaoke), CocoLyricist (AI co-creation for stroke recovery), and SinTechSVS (expressive singing voice synthesis). The latest articles highlight a strong trend in AI-driven music and health technologies , particularly in controllable lyric generation, singing voice synthesis, automatic pronunciation assessment, and multimodal music transcription. The research increasingly integrates large language models, explainable AI, fairness, and real-world deployment, reflecting a shift from theoretical exploration to practical, human-centered applications in healthcare and education. Dr. Wang has received numerous scientific honors, including: Best Paper Awards at ACM MM, ISMIR, IEEE ISM, and CHI First Prize, Asia Pacific Assistive, Rehabilitative, and Therapeutic Technologies Challenge (2015) Faculty Teaching Excellence Award, NUS School of Computing (2024) Top Paper Award, ACM Multimedia 2022 AI in Medicine Collaborative Grant for CocoLyricist project He has supervised over 11 PhD and 20 MComp students and is currently guiding six PhD candidates. His grants come from MOE, NRF, A*STAR, Nokia, and Smule. He has served as General Chair of ISMIR2017 and TPC Co-Chair of ICOT2017, and is on the editorial boards of IEEE Transactions on Multimedia and Journal of New Music Research. He has also developed and taught the first course on Sound and Music Computing in Singapore. Dr. Wang leads the Sound and Music Computing Lab (SMC Lab) , a multidisciplinary team exploring the synergy of music computing, AI, mobile technology, and cloud systems for health and education. The lab actively collaborates with medical institutions such as NUS Yong Loo Lin School of Medicine, Singapore General Hospital, and Harvard Medical School, and is currently working on projects in AI-supported language learning, stroke rehabilitation, and intelligent music interfaces.
Prof. Roger Wattenhofer is a Full Professor at the Department of Information Technology and Electrical Engineering, ETH Zurich, Switzerland, and Deputy head of the Computer Engineering and Networks Lab. He holds a doctorate in Computer Science from ETH Zurich (1998) and has held positions at Brown University and Microsoft Research before returning to ETH. His research focuses on distributed computing, wireless networks, and algorithmic systems design, with contributions to Byzantine agreement protocols, blockchain technologies, and neural network architectures. He teaches courses such as Distributed Systems and Computational Thinking. Education: Ph.D. in Computer Science (ETH Zurich, 1998). Research interests include distributed systems, network algorithms, and the intersection of machine learning with distributed computing. His work spans both theoretical foundations and practical implementations, addressing challenges in fault tolerance, consensus mechanisms, and algorithmic efficiency. Recent publications explore topics like adversarial robustness in voting systems, privacy in reinforcement learning, and generative music models. He actively contributes to open-source frameworks and benchmarks for neural algorithmic reasoning.
Prof. Matthias Nießner is a Professor at the Technical University of Munich , where he leads the Visual Computing Lab . Prior to this, he held a Visiting Assistant Professor position at Stanford University . His work bridges computer vision , graphics , and machine learning , focusing on 3D reconstruction , semantic scene understanding , and AI-driven video synthesis . Prof. Nießner has published over 150 works in top venues like SIGGRAPH , CVPR , and ECCV , with several receiving best paper awards (SIGCHI’14, HPG’15, SPG’18, SIGGRAPH’16 Emerging Tech). His research has garnered international media attention, including features in the New York Times , Wall Street Journal , and MIT Technological Review , as well as TV demonstrations (e.g., Jimmy Kimmel Live for Face2Face technology). Awards : TUM-IAS Rudolph Moessbauer Fellowship (2017–ongoing) Google Faculty Award (2017) Nvidia Professor Partnership Award (2018) ERC Starting Grant (2018, €1.5M) Eurographics Young Researcher Award (2019) Research Trends : 3D Gaussian Splatting for real-time rendering Neural Radiance Fields (NeRF) with mesh supervision Audio-driven facial animation via diffusion models Latent space diffusion for 3D scenes Self-supervised and zero-shot methods for 3D and image analysis As a co-founder and director of Synthesia Inc. , he drives democratization of synthetic media. His YouTube channel has over 5 million views, reflecting his impact beyond academia.
Mark Gales is Professor of Information Engineering at the University of Cambridge and an Official Fellow at Emmanuel College. He is currently on sabbatical leave for the 2024/25 academic year. Prior to his academic career, he worked as a consultant at Roke Manor Research Ltd, developing radar systems, before transitioning to speech and language processing. PhD in 'Model-Based Techniques for Robust Speech Recognition' (University of Cambridge, 1995) BA in Electrical and Information Sciences (University of Cambridge, 1988) His research focuses on speech and language processing , particularly in automated language assessment and low-resource speech technology . He leads the Automated Language Teaching and Assessment (ALTA) Institute , which collaborates with Cambridge University Press & Assessment (CUP&A) to develop commercial tools like Linguaskill and Speak & Improve . These platforms provide automated spoken/written assessment for millions of users globally. Recent publications highlight his work in LLM-driven speech processing , including adversarial attacks on foundation models, end-to-end spoken error correction, and uncertainty estimation frameworks. His team's research spans multilingual capabilities, with deployments in languages ranging from Dholuo to Tok Pisin . Awards : IEEE Fellow, ISCA Fellow Leadership : Fellows' Steward at Emmanuel College Mark has contributed extensively to Hidden Markov Model (HMM) applications in speech recognition, which underpinned early automatic speech systems. His work now bridges LLM-based language assessment with cross-lingual transfer learning and robustness testing for real-world deployments.
Professor Gabriel Brostow is a faculty member in the Department of Computer Science at University College London (UCL), where he leads research in Computer Vision and Human-Computer Interaction. He also serves as Chief Research Scientist and Senior Director of the R&D Team at Niantic, the company behind Pokémon GO. His work bridges academic research and industry applications, focusing on developing AI systems that enhance human capabilities through what he terms 'Human in the Loop AI'—now commonly referred to as Human-Centered AI. Brostow completed his BS in Electrical Engineering at UT Austin, followed by a PhD with Irfan Essa at Georgia Tech. He then pursued postdoctoral research with Roberto Cipolla's Computer Vision & Robotics Group at Cambridge University as a Marshall Sherfield Fellow, and with Marc Pollefeys in ETH Zurich's CVG Group. His research explores how AI, particularly Computer Vision, can serve as 'super-tools' for professionals across various domains including filmmaking, architecture, robotics, and scientific research. Specific interests include assistive technology for everyday life, authoring systems that maximize user effort, 3D reconstruction, depth estimation, and vision-language models. His work often involves creating systems that are validated through real-world human interaction to ensure practical utility. Analysis of his recent publications reveals a strong focus on practical applications of Computer Vision that directly interact with humans. His research spans 3D scene understanding, depth estimation, sketch-based interfaces, and multimodal AI systems. There's a clear emphasis on creating benchmarks and tools that facilitate human-AI collaboration, with applications in assistive technology, urban planning, filmmaking, and biodiversity monitoring. His work frequently appears at top conferences including CVPR, NeurIPS, ECCV, and CHI. Marshall Sherfield Fellowship Brostow actively mentors PhD students, with current advisees including Ross Murphy, Skanda Koppula, Gizem Unlu, Omiros Pantazis, and Jamie Watson. His alumni include numerous PhD graduates and MSc students who have gone on to successful careers in academia and industry. He emphasizes selecting students based on passion and potential rather than just academic credentials, valuing traits like helpfulness, drive, and hunger to learn. His research is supported through collaborations with major institutions and companies including DeepMind, MIT, and the University of Edinburgh. He leads a research group at UCL that collaborates closely with Niantic's R&D team, creating a unique bridge between academic research and industry application. His team's work frequently involves developing novel Computer Vision techniques that are validated through real-world human interaction, ensuring practical utility alongside technical innovation. The group explores blue-sky research problems with applications ranging from assistive technology to professional tools for filmmakers, architects, and scientists studying diverse environments.
Julian McAuley is a Professor in the Department of Computer Science and Engineering at the University of California, San Diego's Jacobs School of Engineering. His research spans recommender systems, machine learning, natural language processing, music information retrieval, and multimodal learning. He maintains an active research group with numerous PhD students and postdocs working on cutting-edge AI problems. His research interests focus on developing advanced algorithms for personalized recommendation systems, with particular emphasis on sequential recommendation, multimodal learning, and integrating large language models with traditional recommendation approaches. His work bridges the gap between theoretical machine learning and practical applications across multiple domains including e-commerce, music, and healthcare. McAuley has published extensively in top-tier conferences including NeurIPS, ICML, KDD, SIGIR, and ACL, with his most recent work exploring the intersection of large language models and recommendation systems. His publications reveal a strong trend toward multimodal approaches that combine text, vision, and audio for more comprehensive understanding and recommendation. He has received significant research funding from major technology companies including Google, Amazon, Facebook, Adobe, and Samsung, as well as government agencies like the National Science Foundation and Department of Defense. His work has practical applications across multiple industries, with a focus on improving user experience through better personalization. McAuley advises numerous PhD students who have gone on to successful careers at leading technology companies and academic institutions. His former students include Wang-Cheng Kang and Jianmo Ni at Google DeepMind, Chris Donahue and Zachary Lipton as assistant professors at CMU, and Ruining He at Google Deepmind.
Dr. George Stamou is a Professor at the School of Electrical and Computer Engineering of the National Technical University of Athens (NTUA), serving as Director of the Artificial Intelligence and Learning Systems Laboratory (AILS). His expertise spans knowledge representation, machine learning, neural networks, and semantic technologies. He leads interdisciplinary initiatives such as the postgraduate program 'Data Science and Machine Learning' (2018–2022). Research Interests: Focuses on knowledge graphs, interpretable AI, semantic web applications, and multimodal learning. His work integrates formal logic systems (e.g., description logics) with modern deep learning techniques, addressing challenges in explainability, bias detection, and ethical AI applications. Publications: Over 150 articles in AI journals/conferences with an h-index of 34 (Google Scholar). Notable contributions include datasets like CHORDONOMICON (music analysis), GOSt-MT (gender bias in MT), and methodologies for counterfactual explanations in machine learning. Awards & Committees: Active in W3C and RuleML standardization bodies. Co-organized major AI conferences. Recognized for contributions to semantic interoperability and knowledge-based systems. Labs & Teams: Directs AILS-NTUA lab and collaborates with CISRI (Computer & Information Systems Research Institute). Engages in EU projects like CultureLabs (cultural heritage digitalization) andsmarty4covid (health data analysis).
Nils Holzenberger is an Assistant Professor at Télécom Paris, France, since February 2023, affiliated with the Data Intelligence Graphs (DIG) research team within the Information Processing and Communication Laboratory (LTcI). His work bridges artificial intelligence, natural language processing, and legal domains through neuro-symbolic approaches to statutory reasoning, particularly in tax law. Education: PhD in Computer Science, Johns Hopkins University (2017-2022) Master's in Engineering, Mines ParisTech (2013-2017) Preparatory Classes, Lycée Louis-le-Grand (2011-2013) Holzenberger's research centers on legal artificial intelligence with emphasis on statutory reasoning limitations in large language models. He pioneered the SARA dataset for tax law reasoning and LegalBench benchmark, developing hybrid symbolic-neural frameworks that expose LLMs' shortcomings in precise legal interpretation. His work integrates Prolog solvers with NLP techniques to create executable tax code mappings and contract analysis tools, establishing foundational methods for verifiable legal AI systems. Analysis of his 15 most recent publications (2019-2024) reveals three dominant research thrusts: (1) Tax law reasoning benchmarks exposing LLM hallucinations, (2) Neuro-symbolic integration for statutory interpretation, and (3) Low-resource template extraction for legal documents. His work consistently demonstrates that pure neural approaches fail at precise legal reasoning, necessitating symbolic grounding for reliable legal AI applications. The DIG research team at Télécom Paris, where Holzenberger leads legal AI initiatives, is actively hiring faculty for neuro-symbolic projects. While specific grant details aren't public, his collaborations with HEC Paris, Copilex startup, and featured podcast appearances indicate substantial industry-academia engagement in legal tech development. Holzenberger directs the legal AI vertical within LTcI's DIG team, focusing on data intelligence for statutory reasoning. His group develops tools for tax minimization strategy discovery, contract analysis, and legal information extraction, maintaining close ties with legal practitioners through projects like the Prolog-based tax code interpreter. The team's infrastructure supports both academic research and startup partnerships in computational law.
Danqi Chen is an Associate Professor in the Department of Computer Science at Princeton University's School of Engineering and Applied Science. Their research focuses on advancing large language models (LLMs), with emphasis on model alignment, safety, and long-context reasoning capabilities. Key research areas: LLMs, AI safety, retrieval systems, and model optimization Recent work explores theorem proving, context encoding, and ethical content generation Their 2025 publications highlight innovations in formal verification scaffolding, attention mechanism efficiency, and copyright-aware generation. 2024 studies investigate continual memorization, rule-based chatbot representations, and scientific literature retrieval benchmarks. Current projects demonstrate commitment to improving model robustness, interpretability, and security compliance in multimodal systems.
David J. Crandall is the Luddy Professor of Computer Science at Indiana University's Luddy School of Informatics, Computing, and Engineering. He serves as Director of the Luddy Artificial Intelligence Center and leads the IU Computer Vision Lab. With joint appointments in Informatics, Cognitive Science, Data Science, and Statistics, his work spans computer vision, machine learning, and AI. He holds a Ph.D. from Cornell University and previously worked at Eastman Kodak Research Labs. His research focuses on developing statistical and machine learning methods to analyze visual information, including object recognition, human activity analysis in video, 3D reconstruction, social media mining, and computational studies of visual attention. Key applications include egocentric vision systems, social robotics for healthcare, and cross-disciplinary collaborations with developmental psychology. Recent publications demonstrate strong emphasis on egocentric video analysis (Ego4D), human-robot interaction (CHI/HRI), and explainable AI (IJCAI). Medical imaging, nanoscale security systems, and computational social science represent emerging interdisciplinary directions. His work consistently integrates deep learning with real-world applications in health, environmental monitoring, and cultural analytics. Tracy M. Sonneborn Award (2024) Distinguished Member of the ACM (2023) Luddy Professorship (2021) NSF CAREER Grant (2013) Trustees Teaching Award (2017) He has advised over 20 Ph.D. graduates, with current students working on computer vision, robotics, and AI ethics. Major grants include $20M for the NSF AI Institute on Engaged Learning, $4.4M for trusted AI research, and funding from NIH, Google, ONR, and NASA. He directs the Computer Vision Lab and collaborates with Selma Sabanović's robotics group on social agents for older adults.
Rozenn Dahyot is a Professor of Computer Science at Maynooth University within the Faculty of Science & Engineering. She previously held roles as Assistant and Associate Professor in Statistics at Trinity College Dublin (2008-2021) and Lecturer in Computer Science (2005-2008). Her research interests bridge Digital Signal Processing, Computer Vision, Machine Learning, and Statistical Analysis. She organized the European Signal Processing Conference (EUSIPCO2021) in Dublin and served as President of the Irish Pattern Recognition and Classification Society (IPRCS) from 2014-2020. Her work spans topics like semantic scene understanding, CNN compression, and medical image segmentation. Key contributions include advancements in graph-based image analysis, reinforcement learning optimization, and AI-driven systems for disaster management. Dahyot is a member of IEEE, ACM, and EURASIP, contributing to both academic and industrial collaborations.
Muchao Ye is an Assistant Professor in the Department of Computer Science at the University of Iowa. He earned his Ph.D. from Pennsylvania State University's College of Information Sciences and Technology in 2024 and a Bachelor of Engineering in Information Engineering from South China University of Technology. Ph.D., Information Sciences and Technology, Pennsylvania State University (2024) B.Eng., Information Engineering, South China University of Technology His research focuses on the intersection of Artificial Intelligence, Machine Learning, and AI Safety, particularly adversarial robustness in language models and vision-language models. He designs methods to enhance the security and reliability of deep learning systems for safety-critical applications like video surveillance and healthcare. Recent publications highlight adversarial robustness frameworks (e.g., UniT , PAT ), vision-language models for explainable video anomaly detection ( VERA ), and healthcare risk prediction techniques ( MedPath , MedRetriever ). His work appears in top venues such as NeurIPS, KDD, AAAI, ACL, and CVPR. Professional experience includes Applied Scientist internships at Amazon (2022–2023) and teaching roles at the University of Iowa and Pennsylvania State University. He serves as a reviewer for conferences like NeurIPS, ICML, and journals including IEEE TPAMI.
Noah A. Smith is an Adjunct Professor of Computer Science and Engineering at the University of Washington. His work focuses on computational linguistics, machine learning, and natural language processing. He holds a Ph.D. in Computer Science from Johns Hopkins University (2006). His research explores ethical AI applications, multimodal systems, and foundational aspects of language models. Key research areas include: Ethical considerations in NLP, such as detecting rights abuses through text analysis Efficient decoding and alignment strategies for large language models Large-scale evaluation frameworks for multitask and multimodal generation Understanding pretraining dynamics and data composition effects Recent work emphasizes transparency in language models (e.g., tracing outputs to training data) and improving alignment through human feedback. He has contributed to open-source projects like OLMo and Dolma, advancing reproducibility in NLP research. No awards explicitly listed in provided texts. No specific advising or grant details available, though extensive publication output indicates active research involvement.