Almut Sophia Koepke is a junior research group leader at the Technical University of Munich and University of Tübingen, focusing on multimodal learning problems integrating sound, vision, and text. Her work bridges foundational research in audio-visual understanding with practical applications in few-shot learning, zero-shot translation, and cross-modal attention mechanisms.
Oliver Kroemer is an Associate Professor at Carnegie Mellon University's Robotics Institute (RI), affiliated with the Intelligent Autonomous Manipulation (IAM) Lab. His research focuses on enabling robots to learn versatile manipulation skills through lifelong frameworks, with applications in elder care, environmental maintenance, and hazardous operations. Developed methods for robot learning via physical interaction and reinforcement learning Created representations for contact states and motor primitives to improve skill generalization Research Interests: Spanning robot learning, tactile sensing, force-velocity control, and lifelong skill acquisition. Projects include Agile and Dynamic Interactions for Mobile Manipulation and Integrated Planning and Learning (Pillar project). Scientific Awards: Finalist, Georges Giralt Ph.D. Award (2015) Education: Masters & Bachelors in Engineering, University of Cambridge (2008) Ph.D., Technische Universitaet Darmstadt (2014) Students & Affiliates: Current PhD: Mark Lee, Sarvesh Patil, Saumya Saxena, Yunus Seker, Zilin Si Past PhD: Alex LaGrassa, Tabitha Lee, Qiao Liang, Shivam Vats, Kevin Zhang
Timothy M. Hospedales is a Professor of Artificial Intelligence at the Institute of Perception, Action and Behaviour within the School of Informatics at the University of Edinburgh . He also serves as VP AI and Head of Samsung AI Research Centre Europe . His research focuses on efficient and robust AI , emphasizing meta-learning , lifelong transfer-learning , and domain adaptation in both probabilistic and deep learning frameworks. Applications span computer vision , vision and language , reinforcement learning for robotics , and finance . Professor at University of Edinburgh (2020–present) ELLIS Fellow (2021) Head of Samsung AI Research Europe (2020–present) Founding Director of Applied Machine Learning Lab at QMUL (2012–2016) His work includes pioneering contributions to meta-learning , few-shot learning , and self-supervised methods , with notable awards such as the Best Paper Prize at ICML AutoML 2018 and Best Student Paper at ICPR 2018 . He has co-authored 15+ recent papers on topics like Vision-Language Models , Medical AI Fairness , and Diffusion Model Optimization . He served as Program Co-Chair for BMVC 2018 and AAAI 2022 , and authored a book on Visual Adaptation in the Deep Learning Era (2022). Co-Chair, BMVC 2018 Guest Editor, IET CV Special Issue (2016) Keynote Speaker at TASK-CV Workshop (ECCV 2016) Special Issue on Fewer Labels (IEEE PAMI 2020) His leadership extends to organizing workshops like the Learning-to-Learn Workshop at ICLR 2021 , Meta-Learning Workshop at NeurIPS 2020 , and Domain Generalisation Workshop at ICLR 2023 . Current projects include Meta-Omnium (CVPR 2023) for general-purpose meta-learning and MetaAudio (ICANN 2022) for few-shot audio classification benchmarks.
Professor Carlo Harvey is a creative technologist at the School of Digital Arts (SODA), Manchester Metropolitan University. His interdisciplinary research merges games , machine learning , virtual production , and cultural heritage reinterpretation . He leads industry collaborations with entities like Jaguar Land Rover and Epic Games, focusing on AI-driven interactive audio, real-time visualization, and accessibility solutions. Award-winning projects : TIGA, Innovate UK, and Epic Games MegaGrant for Accession Industry partnerships : Automotive sector, cultural institutions His research spans human-computer interaction , multisensory virtual environments , and acoustic-visual cross-modal perception . Recent publications address robotic simulations, motion alignment, and haptic feedback systems. Scientific recognition : TIGA Award, Innovate UK Funding, Epic Games MegaGrant Advocacy : Digital inclusion, creative collaboration, social impact of technology
University of Illinois Urbana-ChampaignUnited States
Pasquale Bottalico serves as Associate Professor in the Department of Speech and Hearing Science at the University of Illinois, with dual appointments as Associate Professor at the Center for Latin American and Caribbean Studies and Affiliate Faculty in the School of Music. His unique interdisciplinary profile bridges engineering, music performance, and speech science, reflecting his dual academic training and professional artistry. His educational foundation includes: Bachelor's in Telecommunications Engineering from Univeristà Mediterranea di Reggio Calabria, Italy Concurrent Opera Singing degree from F. Cilea Music Academy, Reggio Calabria Master's in Telecommunications Engineering from Politecnico di Torino, Italy Ph.D. in Metrology specializing in acoustics measurement uncertainty and classroom acoustics Dr. Bottalico's research centers on vocal load quantification and professional voice techniques , with significant contributions to understanding vocal fatigue in teachers and singers. His work spans Speech Intelligibility in educational environments, Room Acoustics for performance and learning spaces, and Musical Acoustics of historical vocal styles. A distinctive thread throughout his research examines how acoustic conditions modulate voice production and perception, increasingly incorporating virtual reality and bone conduction technologies for innovative assessment and intervention approaches. His Colombian vocal health study demonstrates cross-cultural applications of his work. Analysis of his 2023-2025 publications reveals three dominant research trajectories: (1) The impact of noise and dysphonia on children's speech processing in educational settings, using multimodal assessment including EEG; (2) Virtual reality applications for voice production research and therapeutic intervention; (3) Cross-cultural validation of vocal fatigue metrics and development of biofeedback systems. His work consistently bridges engineering precision with clinical applicability, particularly for professional voice users in challenging acoustic environments. No scientific awards were documented in the available information. While specific advising relationships aren't detailed, his research collaborations span international institutions including Colombian and Italian universities, suggesting graduate mentorship in interdisciplinary projects. No grant information was provided, though his systematic reviews and cross-cultural studies imply externally funded research activities. Though no dedicated laboratory is specified, his virtual reality voice studies and acoustic parameter assessments suggest affiliations with audio engineering facilities and voice clinics, likely through the Speech and Hearing Science department's research infrastructure.
Prof. Matthias Nießner is a Professor at the Technical University of Munich , where he leads the Visual Computing Lab . Prior to this, he held a Visiting Assistant Professor position at Stanford University . His work bridges computer vision , graphics , and machine learning , focusing on 3D reconstruction , semantic scene understanding , and AI-driven video synthesis . Prof. Nießner has published over 150 works in top venues like SIGGRAPH , CVPR , and ECCV , with several receiving best paper awards (SIGCHI’14, HPG’15, SPG’18, SIGGRAPH’16 Emerging Tech). His research has garnered international media attention, including features in the New York Times , Wall Street Journal , and MIT Technological Review , as well as TV demonstrations (e.g., Jimmy Kimmel Live for Face2Face technology). Awards : TUM-IAS Rudolph Moessbauer Fellowship (2017–ongoing) Google Faculty Award (2017) Nvidia Professor Partnership Award (2018) ERC Starting Grant (2018, €1.5M) Eurographics Young Researcher Award (2019) Research Trends : 3D Gaussian Splatting for real-time rendering Neural Radiance Fields (NeRF) with mesh supervision Audio-driven facial animation via diffusion models Latent space diffusion for 3D scenes Self-supervised and zero-shot methods for 3D and image analysis As a co-founder and director of Synthesia Inc. , he drives democratization of synthetic media. His YouTube channel has over 5 million views, reflecting his impact beyond academia.
Xavier Serra is a Full Professor at the Department of Engineering at Universitat Pompeu Fabra (UPF), Barcelona. He is the founder and director of the Music Technology Group (MTG), and leads the UPF-BMAT Chair on AI and Music. He also coordinates the Master in Sound and Music Computing and serves as President of the Phonos Foundation. His research focuses on audio signal processing, sound and music computing, and computational musicology, emphasizing open science and open innovation. Education: BSc in Biology, University of Barcelona (1981) Master in Music, Florida State University (1983) PhD in Computer Music, Stanford University (1989) Research Interests: Audio Signal Processing Data-Driven and Knowledge-Driven Methodologies Music Information Retrieval Cultural Music Analysis (e.g., Carnatic/Turkish/Andalusian Music) Music Education Technology Notable Projects: CompMusic (ERC Advanced Grant, 2010-2017): Multicultural computational music analysis Open datasets: Freesound, Saraga, FSD50K Technologies: Reactable, Vocaloid, Essentia API Recent Trends in Articles: Focus on AI-driven audio processing (neural fingerprints, generative models), cross-cultural music analysis, and explainable music difficulty estimation. Awards: ERC Advanced Grant (2010) for CompMusic Project. Labs/Teams: Director of MTG, Phonos Foundation, and UPF-BMAT Chair. Active in open-source projects and international collaborations.
David W. Jacobs is a Professor in the Department of Computer Science at the University of Maryland, with a joint appointment at the University of Maryland Institute for Advanced Computer Studies (UMIACS). He also served as the interim Director of the University of Maryland Center for Machine Learning starting in 2018. University: University of Maryland School: College of Computer, Mathematical, and Natural Sciences Department: Department of Computer Science Academic Rank: Professor Education: He received his B.A. from Yale University, and M.S. and Ph.D. in Computer Science from MIT. Research Interests: His research primarily focuses on computer vision and machine learning, particularly visual object recognition, lighting variation modeling, 3D reconstruction, perceptual organization, motion understanding, and the integration of vision with graphics and human-computer interaction. A major applied contribution is the development of Leafsnap , an electronic field guide app for plant identification, which has been downloaded over 1.5 million times and used in biodiversity and educational contexts. Publication Trends: His recent scholarly output centers on deep learning, convolutional networks, residual architectures, generative models (especially GANs), and interpretability. His work often bridges theoretical insights with practical applications in vision and AI. Scientific Awards: Honorable Mention, Best Paper Award, CVPR 2000 Best Student Paper Award, UIST 2003 Best Paper Award, Eurographics 2016 2011 Edward O. Wilson Biodiversity Technology Pioneer Award for Leafsnap Teaching and Advising: He has taught advanced courses such as CMSC 422 (Introduction to Machine Learning) and CMSC 828L (Deep Learning). He mentors students through course projects and research, though specific advisees are not listed. He has collaborated with institutions like Columbia University and the Smithsonian on impactful interdisciplinary projects. Labs and Teams: He is affiliated with UMIACS and leads research efforts in vision and learning, contributing to the University of Maryland Center for Machine Learning. His team has developed several mobile applications including Leafsnap, Birdsnap, and Dogsnap, demonstrating a strong focus on real-world deployment of vision technology.
University of California , Santa Barbara (UCSB)United States
Yao Qin is an Assistant Professor in the Department of Electrical and Computer Engineering at the University of California, Santa Barbara (UCSB), with dual affiliation in the Department of Computer Science. She concurrently serves as Co-Director of the REAL AI Initiative at UCSB and holds a Senior Research Scientist position at Google DeepMind, where she contributes to the Gemini Multimodal project. Her academic credentials include a PhD in Computer Science and Engineering from the University of California, San Diego (advised by Prof. Garrison W. Cottrell) and a BS in Electrical Engineering from Dalian University of Technology. During her doctoral studies, she completed internships with pioneering researchers Geoffrey Hinton and Ian Goodfellow. Dr. Qin's research program centers on machine learning robustness, with emphasis on adversarial robustness, out-of-distribution generalization, and fairness. She develops reliable AI systems specifically for healthcare applications, with diabetes management as a primary focus. Her lab explores critical themes including AI safety in multimodal models and diabetes-specific AI solutions, particularly exercise metabolism modeling and glycemic effect prediction. Recent publications reveal a strong trajectory in robust machine learning with cross-domain applications. Her work consistently bridges theoretical robustness concepts with practical healthcare implementations, particularly in diabetes care. Key publication venues include CVPR, ICML, NeurIPS, and ICLR, with notable contributions to out-of-distribution detection, adversarial transfer learning, and multimodal AI safety. Her distinguished recognition includes: EECS Rising Star at MIT (2021) UCSB Regents' Junior Faculty Fellowship Award Helmsley Charitable Trust award for Type 1 diabetes research UCSB Faculty Research Grant American Diabetes Association Abstract Award (ADA-2025) Dr. Qin actively mentors four PhD students—Mehak Dhaliwal, Andong Hua, Kenan Tang, and Youngseok Yoon—on projects spanning LLMs for diabetes, multimodal robustness, and generative time-series modeling. Her research is funded by the Helmsley Charitable Trust and UCSB, with recent grants supporting exercise-specific AID algorithms for diabetes management. As Co-Director of the REAL AI Initiative, she leads a research ecosystem focused on developing reliable artificial intelligence. Current lab activities include organizing workshops at NeurIPS-2024 (AdvML-Frontiers and AIM-FM) and developing next-generation diabetes management tools through collaborations with medical institutions.
Takuya Tsunoda serves as Assistant Professor of Japanese Film and Media in the Department of East Asian Languages & Cultures at Columbia University, with prior teaching appointments at Colgate University and the University of Chicago. Currently on leave for spring 2022, he maintains his office in 416 Kent Hall (contact: (212) 854-5040, tt2101@columbia.edu). His academic credentials include: BA: Waseda University (2002), Columbia University (2005) MA: Columbia University (2008) PhD: Yale University (2015) Dr. Tsunoda's research centers on institutional-media interplay through historical and theoretical lenses, examining how technologies shape socio-cultural practices and knowledge formations. His seminal work investigates Iwanami Productions' evolution from educational/science film provider to New Cinema pioneer, arguing that Japan's 1960s cinematic revolution originated in institutionalized audio-visual pedagogy rather than political radicalism. Current explorations extend to media reflexivity, alpine photography, insect ecology, children's media engagement, and diegesis in contemporary visual culture. Teaching : Courses include East Asian Cinema, Japanese Contemporary Media Culture, Japanese New Wave Modernism, Documentary Film Studies, and Critical Approaches to East Asian Theory. His publication trajectory (2018-2022) reveals consistent focus on postwar Japanese media archaeology, connecting educational film, industrial cinema, and New Wave movements through archival analysis. Key themes include cross-medial academicism, mediascape transformations, and the socio-political dimensions of visual pedagogy in reconstructing cinematic modernism. Active research trajectories include transnational New Wave parameters, television documentary studies, and ecological media representations, with ongoing book project examining Iwanami Productions' institutional legacy.
Nima Mesgarani is an Associate Professor of Electrical Engineering at Columbia Engineering, Columbia University, affiliated with the Sense, Collect and Move Data Committee. His research bridges engineering and neuroscience through reverse-engineering neural signal processing mechanisms, leading to advancements in brain-machine interfaces, neural prosthetics, and speech processing algorithms. He received his PhD in Electrical Engineering from the University of Maryland and completed postdoctoral training at Johns Hopkins University's Center for Language and Speech Processing and UC San Francisco's Neurosurgery Department. Research Focus Professor Mesgarani's lab integrates computational neuroscience and engineering to study acoustic signal processing. Key areas include: Neural decoding of speech and auditory attention in multi-talker environments Development of brain-controlled hearing technologies Novel speech separation and synthesis algorithms inspired by cortical processing Cross-modal learning between auditory and visual systems Applications of large language models in neural signal interpretation Publication Trends Analysis of his 15 most recent articles (2025) reveals dominant themes: neural decoding techniques using intracranial EEG, brain-inspired speech separation models (e.g., Mamba architectures), applications of large language models in auditory neuroscience, cross-modal distillation methods, and clinical translation of audio processing algorithms. A strong emphasis emerges on real-time brain-computer interfaces and noise-robust speech processing. Laboratory and Collaborations Mesgarani directs an interdisciplinary lab developing neurotechnology for hearing restoration. His team collaborates with neurosurgery departments and speech processing centers, focusing on translating theoretical models into clinical brain-machine interfaces. The lab's work has yielded patents for brain-informed speech separation systems and attention-decoding frameworks.
Professor Gabriel Brostow is a faculty member in the Department of Computer Science at University College London (UCL), where he leads research in Computer Vision and Human-Computer Interaction. He also serves as Chief Research Scientist and Senior Director of the R&D Team at Niantic, the company behind Pokémon GO. His work bridges academic research and industry applications, focusing on developing AI systems that enhance human capabilities through what he terms 'Human in the Loop AI'—now commonly referred to as Human-Centered AI. Brostow completed his BS in Electrical Engineering at UT Austin, followed by a PhD with Irfan Essa at Georgia Tech. He then pursued postdoctoral research with Roberto Cipolla's Computer Vision & Robotics Group at Cambridge University as a Marshall Sherfield Fellow, and with Marc Pollefeys in ETH Zurich's CVG Group. His research explores how AI, particularly Computer Vision, can serve as 'super-tools' for professionals across various domains including filmmaking, architecture, robotics, and scientific research. Specific interests include assistive technology for everyday life, authoring systems that maximize user effort, 3D reconstruction, depth estimation, and vision-language models. His work often involves creating systems that are validated through real-world human interaction to ensure practical utility. Analysis of his recent publications reveals a strong focus on practical applications of Computer Vision that directly interact with humans. His research spans 3D scene understanding, depth estimation, sketch-based interfaces, and multimodal AI systems. There's a clear emphasis on creating benchmarks and tools that facilitate human-AI collaboration, with applications in assistive technology, urban planning, filmmaking, and biodiversity monitoring. His work frequently appears at top conferences including CVPR, NeurIPS, ECCV, and CHI. Marshall Sherfield Fellowship Brostow actively mentors PhD students, with current advisees including Ross Murphy, Skanda Koppula, Gizem Unlu, Omiros Pantazis, and Jamie Watson. His alumni include numerous PhD graduates and MSc students who have gone on to successful careers in academia and industry. He emphasizes selecting students based on passion and potential rather than just academic credentials, valuing traits like helpfulness, drive, and hunger to learn. His research is supported through collaborations with major institutions and companies including DeepMind, MIT, and the University of Edinburgh. He leads a research group at UCL that collaborates closely with Niantic's R&D team, creating a unique bridge between academic research and industry application. His team's work frequently involves developing novel Computer Vision techniques that are validated through real-world human interaction, ensuring practical utility alongside technical innovation. The group explores blue-sky research problems with applications ranging from assistive technology to professional tools for filmmakers, architects, and scientists studying diverse environments.
Dhruv Jain is an Assistant Professor in the Computer Science and Engineering Department at the University of Michigan , with affiliations in the School of Information and Michigan Medicine . He leads the Soundability Lab , focusing on transforming hearing into a programmable interface through human-centered AI systems. PhD from University of Washington MS from MIT Media Lab Formerly worked at Microsoft Research, Google, and Apple Research Interests intersect Human-Computer Interaction (HCI) , Audio AI , Accessibility , and Hearing Health . His work develops systems like SoundWatch (smartwatch-based sound awareness) and HomeSound (IoT sound visualization) to expand auditory access for DHH individuals and other domains. Recent Publications (2023–2025) span top venues like CHI , ASSETS , and ICMI , with trends including generative AI for real-time captioning , adaptive soundscapes , and clinical communication tools . SIGCHI Outstanding Dissertation Award (2023) William Chan Memorial Dissertation Award (2023) Best Poster Award at ASSETS 2024 Google Academic Research Award (2024) NIH Grant ($450k, 2023) William Demant Foundation Grant ($740k, 2025) Teaching includes Accessible Computing (undergraduate) and Advanced Accessibility (graduate), alongside global DIY workshops in five countries. The Soundability Lab collaborates with neuroscientists, clinicians, and Deaf scholars to create deployable systems like clinical communication tools now at Michigan Medicine and features integrated into Apple devices.
Schloss Dagstuhl - Leibniz Center for InformaticsGermany
Julian McAuley is a Professor in the Department of Computer Science and Engineering at the University of California, San Diego's Jacobs School of Engineering. His research spans recommender systems, machine learning, natural language processing, music information retrieval, and multimodal learning. He maintains an active research group with numerous PhD students and postdocs working on cutting-edge AI problems. His research interests focus on developing advanced algorithms for personalized recommendation systems, with particular emphasis on sequential recommendation, multimodal learning, and integrating large language models with traditional recommendation approaches. His work bridges the gap between theoretical machine learning and practical applications across multiple domains including e-commerce, music, and healthcare. McAuley has published extensively in top-tier conferences including NeurIPS, ICML, KDD, SIGIR, and ACL, with his most recent work exploring the intersection of large language models and recommendation systems. His publications reveal a strong trend toward multimodal approaches that combine text, vision, and audio for more comprehensive understanding and recommendation. He has received significant research funding from major technology companies including Google, Amazon, Facebook, Adobe, and Samsung, as well as government agencies like the National Science Foundation and Department of Defense. His work has practical applications across multiple industries, with a focus on improving user experience through better personalization. McAuley advises numerous PhD students who have gone on to successful careers at leading technology companies and academic institutions. His former students include Wang-Cheng Kang and Jianmo Ni at Google DeepMind, Chris Donahue and Zachary Lipton as assistant professors at CMU, and Ruining He at Google Deepmind.
Yonatan Bisk is an Assistant Professor at Carnegie Mellon University (CMU) in the School of Computer Science , with dual appointments in the Language Technologies Institute and Robotics Institute . His research bridges Natural Language Processing (NLP) with robotics, focusing on grounded and embodied language understanding. Assistant Professor, Language Technologies Institute, CMU (2021–Present) Courtesy Appointment, Robotics Institute, CMU Research Themes : Language as a social codification of embodied experience Interpretable multimodal model training Human-robot collaboration frameworks Embodied question-answering systems Selected Trends : His recent publications show increasing focus on cross-modal attention mechanisms (Vid2Robot), error detection in toolchains (Tools Fail), and theory-of-mind reasoning in language agents (SOTOPIA). Multimodal integration spans vision, audio, and robotic control contexts (ANAVI). Labs & Collaborations : Founder of CLAW Lab (Connecting Language to Action and the World) Collaborations with Microsoft Research, Meta Inc, and CMU's REAL (Robotics, Embodied AI, Learning) community