Xavier Serra is a Full Professor at the Department of Engineering at Universitat Pompeu Fabra (UPF), Barcelona. He is the founder and director of the Music Technology Group (MTG), and leads the UPF-BMAT Chair on AI and Music. He also coordinates the Master in Sound and Music Computing and serves as President of the Phonos Foundation. His research focuses on audio signal processing, sound and music computing, and computational musicology, emphasizing open science and open innovation. Education: BSc in Biology, University of Barcelona (1981) Master in Music, Florida State University (1983) PhD in Computer Music, Stanford University (1989) Research Interests: Audio Signal Processing Data-Driven and Knowledge-Driven Methodologies Music Information Retrieval Cultural Music Analysis (e.g., Carnatic/Turkish/Andalusian Music) Music Education Technology Notable Projects: CompMusic (ERC Advanced Grant, 2010-2017): Multicultural computational music analysis Open datasets: Freesound, Saraga, FSD50K Technologies: Reactable, Vocaloid, Essentia API Recent Trends in Articles: Focus on AI-driven audio processing (neural fingerprints, generative models), cross-cultural music analysis, and explainable music difficulty estimation. Awards: ERC Advanced Grant (2010) for CompMusic Project. Labs/Teams: Director of MTG, Phonos Foundation, and UPF-BMAT Chair. Active in open-source projects and international collaborations.
Elena Maria Baralis is a Full Professor at the Department of Control and Computer Science (DAUIN) at the Polytechnic University of Turin. She serves as Pro-Rector, member of the Board of Directors (without voting rights), member of the Academic Senate (without voting rights), and coordinator of the University's Permanent Observatory for Monitoring the Academic Sector. She chairs the Control and Computer Engineering Department and previously chaired the Computer Engineering School from October 2012 to October 2018. Her research interests focus on database systems and data mining, specifically explainable AI, bias detection in data analytics, and machine learning algorithms for big data. Her work spans various application domains including predictive maintenance, Industry 4.0, and healthcare. Recent publications demonstrate her expertise in speech processing, bias mitigation, and innovative neural network architectures like Kolmogorov-Arnold Networks. Her research output shows a clear trend toward addressing fairness and explainability in AI systems while exploring novel approaches to speech and language understanding. Professor Baralis has received significant recognition including becoming a Fellow of the Academy of Sciences of Turin in 2017. She has served as Editor-in-Chief for IEEE Internet of Things Journal (2016-2019) and Knowledge and Information Systems (2014-present). She actively mentors doctoral students including Claudio Savelli (researching Machine Unlearning), Eleonora Poeta, Giuseppe Gallipoli, Alkis Koudounas, and others. Her research is supported by numerous projects including AI4CTI (Artificial Intelligence for Cyber Threat Intelligence, 2025-2028), Smart manufacturing driven by Machine Learning in Industry 4.0 (2019-2020), and I-REACT (2016-2019).
David W. Jacobs is a Professor in the Department of Computer Science at the University of Maryland, with a joint appointment at the University of Maryland Institute for Advanced Computer Studies (UMIACS). He also served as the interim Director of the University of Maryland Center for Machine Learning starting in 2018. University: University of Maryland School: College of Computer, Mathematical, and Natural Sciences Department: Department of Computer Science Academic Rank: Professor Education: He received his B.A. from Yale University, and M.S. and Ph.D. in Computer Science from MIT. Research Interests: His research primarily focuses on computer vision and machine learning, particularly visual object recognition, lighting variation modeling, 3D reconstruction, perceptual organization, motion understanding, and the integration of vision with graphics and human-computer interaction. A major applied contribution is the development of Leafsnap , an electronic field guide app for plant identification, which has been downloaded over 1.5 million times and used in biodiversity and educational contexts. Publication Trends: His recent scholarly output centers on deep learning, convolutional networks, residual architectures, generative models (especially GANs), and interpretability. His work often bridges theoretical insights with practical applications in vision and AI. Scientific Awards: Honorable Mention, Best Paper Award, CVPR 2000 Best Student Paper Award, UIST 2003 Best Paper Award, Eurographics 2016 2011 Edward O. Wilson Biodiversity Technology Pioneer Award for Leafsnap Teaching and Advising: He has taught advanced courses such as CMSC 422 (Introduction to Machine Learning) and CMSC 828L (Deep Learning). He mentors students through course projects and research, though specific advisees are not listed. He has collaborated with institutions like Columbia University and the Smithsonian on impactful interdisciplinary projects. Labs and Teams: He is affiliated with UMIACS and leads research efforts in vision and learning, contributing to the University of Maryland Center for Machine Learning. His team has developed several mobile applications including Leafsnap, Birdsnap, and Dogsnap, demonstrating a strong focus on real-world deployment of vision technology.
Nima Mesgarani is an Associate Professor of Electrical Engineering at Columbia Engineering, Columbia University, affiliated with the Sense, Collect and Move Data Committee. His research bridges engineering and neuroscience through reverse-engineering neural signal processing mechanisms, leading to advancements in brain-machine interfaces, neural prosthetics, and speech processing algorithms. He received his PhD in Electrical Engineering from the University of Maryland and completed postdoctoral training at Johns Hopkins University's Center for Language and Speech Processing and UC San Francisco's Neurosurgery Department. Research Focus Professor Mesgarani's lab integrates computational neuroscience and engineering to study acoustic signal processing. Key areas include: Neural decoding of speech and auditory attention in multi-talker environments Development of brain-controlled hearing technologies Novel speech separation and synthesis algorithms inspired by cortical processing Cross-modal learning between auditory and visual systems Applications of large language models in neural signal interpretation Publication Trends Analysis of his 15 most recent articles (2025) reveals dominant themes: neural decoding techniques using intracranial EEG, brain-inspired speech separation models (e.g., Mamba architectures), applications of large language models in auditory neuroscience, cross-modal distillation methods, and clinical translation of audio processing algorithms. A strong emphasis emerges on real-time brain-computer interfaces and noise-robust speech processing. Laboratory and Collaborations Mesgarani directs an interdisciplinary lab developing neurotechnology for hearing restoration. His team collaborates with neurosurgery departments and speech processing centers, focusing on translating theoretical models into clinical brain-machine interfaces. The lab's work has yielded patents for brain-informed speech separation systems and attention-decoding frameworks.
Professor Gabriel Brostow is a faculty member in the Department of Computer Science at University College London (UCL), where he leads research in Computer Vision and Human-Computer Interaction. He also serves as Chief Research Scientist and Senior Director of the R&D Team at Niantic, the company behind Pokémon GO. His work bridges academic research and industry applications, focusing on developing AI systems that enhance human capabilities through what he terms 'Human in the Loop AI'—now commonly referred to as Human-Centered AI. Brostow completed his BS in Electrical Engineering at UT Austin, followed by a PhD with Irfan Essa at Georgia Tech. He then pursued postdoctoral research with Roberto Cipolla's Computer Vision & Robotics Group at Cambridge University as a Marshall Sherfield Fellow, and with Marc Pollefeys in ETH Zurich's CVG Group. His research explores how AI, particularly Computer Vision, can serve as 'super-tools' for professionals across various domains including filmmaking, architecture, robotics, and scientific research. Specific interests include assistive technology for everyday life, authoring systems that maximize user effort, 3D reconstruction, depth estimation, and vision-language models. His work often involves creating systems that are validated through real-world human interaction to ensure practical utility. Analysis of his recent publications reveals a strong focus on practical applications of Computer Vision that directly interact with humans. His research spans 3D scene understanding, depth estimation, sketch-based interfaces, and multimodal AI systems. There's a clear emphasis on creating benchmarks and tools that facilitate human-AI collaboration, with applications in assistive technology, urban planning, filmmaking, and biodiversity monitoring. His work frequently appears at top conferences including CVPR, NeurIPS, ECCV, and CHI. Marshall Sherfield Fellowship Brostow actively mentors PhD students, with current advisees including Ross Murphy, Skanda Koppula, Gizem Unlu, Omiros Pantazis, and Jamie Watson. His alumni include numerous PhD graduates and MSc students who have gone on to successful careers in academia and industry. He emphasizes selecting students based on passion and potential rather than just academic credentials, valuing traits like helpfulness, drive, and hunger to learn. His research is supported through collaborations with major institutions and companies including DeepMind, MIT, and the University of Edinburgh. He leads a research group at UCL that collaborates closely with Niantic's R&D team, creating a unique bridge between academic research and industry application. His team's work frequently involves developing novel Computer Vision techniques that are validated through real-world human interaction, ensuring practical utility alongside technical innovation. The group explores blue-sky research problems with applications ranging from assistive technology to professional tools for filmmakers, architects, and scientists studying diverse environments.
Dhruv Jain is an Assistant Professor in the Computer Science and Engineering Department at the University of Michigan , with affiliations in the School of Information and Michigan Medicine . He leads the Soundability Lab , focusing on transforming hearing into a programmable interface through human-centered AI systems. PhD from University of Washington MS from MIT Media Lab Formerly worked at Microsoft Research, Google, and Apple Research Interests intersect Human-Computer Interaction (HCI) , Audio AI , Accessibility , and Hearing Health . His work develops systems like SoundWatch (smartwatch-based sound awareness) and HomeSound (IoT sound visualization) to expand auditory access for DHH individuals and other domains. Recent Publications (2023–2025) span top venues like CHI , ASSETS , and ICMI , with trends including generative AI for real-time captioning , adaptive soundscapes , and clinical communication tools . SIGCHI Outstanding Dissertation Award (2023) William Chan Memorial Dissertation Award (2023) Best Poster Award at ASSETS 2024 Google Academic Research Award (2024) NIH Grant ($450k, 2023) William Demant Foundation Grant ($740k, 2025) Teaching includes Accessible Computing (undergraduate) and Advanced Accessibility (graduate), alongside global DIY workshops in five countries. The Soundability Lab collaborates with neuroscientists, clinicians, and Deaf scholars to create deployable systems like clinical communication tools now at Michigan Medicine and features integrated into Apple devices.
Julian McAuley is a Professor in the Department of Computer Science and Engineering at the University of California, San Diego's Jacobs School of Engineering. His research spans recommender systems, machine learning, natural language processing, music information retrieval, and multimodal learning. He maintains an active research group with numerous PhD students and postdocs working on cutting-edge AI problems. His research interests focus on developing advanced algorithms for personalized recommendation systems, with particular emphasis on sequential recommendation, multimodal learning, and integrating large language models with traditional recommendation approaches. His work bridges the gap between theoretical machine learning and practical applications across multiple domains including e-commerce, music, and healthcare. McAuley has published extensively in top-tier conferences including NeurIPS, ICML, KDD, SIGIR, and ACL, with his most recent work exploring the intersection of large language models and recommendation systems. His publications reveal a strong trend toward multimodal approaches that combine text, vision, and audio for more comprehensive understanding and recommendation. He has received significant research funding from major technology companies including Google, Amazon, Facebook, Adobe, and Samsung, as well as government agencies like the National Science Foundation and Department of Defense. His work has practical applications across multiple industries, with a focus on improving user experience through better personalization. McAuley advises numerous PhD students who have gone on to successful careers at leading technology companies and academic institutions. His former students include Wang-Cheng Kang and Jianmo Ni at Google DeepMind, Chris Donahue and Zachary Lipton as assistant professors at CMU, and Ruining He at Google Deepmind.
Takako Fujioka is an Associate Professor of Music at Stanford University, affiliated with the Center for Computer Research in Music and Acoustics (CCRMA). Her research focuses on the neural mechanisms underlying auditory perception, auditory-motor coupling, and music-supported therapy for neurorehabilitation. She holds a Ph.D. in Physiology from the Graduate University for Advanced Studies, Japan, and M.Sc./B.Eng. degrees in Electrical Engineering from Waseda University. Her work combines neurophysiological techniques such as MEG and EEG to study brain plasticity in development, aging, and stroke recovery. Notable contributions include investigating how music influences motor and cognitive recovery in stroke patients, as well as exploring the neural basis of musical perception through rhythmic synchronization and pitch discrimination studies. Supported by awards from the Canadian Institutes of Health Research during her postdoctoral work at the Rotman Research Institute, her research bridges clinical neuroscience and music cognition. Dr. Fujioka’s expertise spans auditory neuroscience, neurorehabilitation, and technology-assisted music therapy. She has pioneered studies on tactile mapping for cochlear implant users and networked music performance systems, emphasizing cross-modal perception and human-technology interaction. Her findings contribute to both theoretical understanding of auditory processing and practical applications in medical and educational settings. Awards: Canadian Institutes of Health Research Awards (postdoctoral phase) Labs/Teams: CCRMA, Stanford Music Perception Laboratory, Rotman Research Institute collaborations Key Themes: Neuroplasticity, Music-Mediated Rehabilitation, Auditory-Motor Integration, Multisensory Processing Her recent work examines aging-related changes in binaural hearing and the role of beta/gamma oscillations in rhythmic processing. She advocates for translational research that connects neural mechanisms with real-world therapeutic interventions.
Dr. George Stamou is a Professor at the School of Electrical and Computer Engineering of the National Technical University of Athens (NTUA), serving as Director of the Artificial Intelligence and Learning Systems Laboratory (AILS). His expertise spans knowledge representation, machine learning, neural networks, and semantic technologies. He leads interdisciplinary initiatives such as the postgraduate program 'Data Science and Machine Learning' (2018–2022). Research Interests: Focuses on knowledge graphs, interpretable AI, semantic web applications, and multimodal learning. His work integrates formal logic systems (e.g., description logics) with modern deep learning techniques, addressing challenges in explainability, bias detection, and ethical AI applications. Publications: Over 150 articles in AI journals/conferences with an h-index of 34 (Google Scholar). Notable contributions include datasets like CHORDONOMICON (music analysis), GOSt-MT (gender bias in MT), and methodologies for counterfactual explanations in machine learning. Awards & Committees: Active in W3C and RuleML standardization bodies. Co-organized major AI conferences. Recognized for contributions to semantic interoperability and knowledge-based systems. Labs & Teams: Directs AILS-NTUA lab and collaborates with CISRI (Computer & Information Systems Research Institute). Engages in EU projects like CultureLabs (cultural heritage digitalization) andsmarty4covid (health data analysis).
Yonatan Bisk is an Assistant Professor at Carnegie Mellon University (CMU) in the School of Computer Science , with dual appointments in the Language Technologies Institute and Robotics Institute . His research bridges Natural Language Processing (NLP) with robotics, focusing on grounded and embodied language understanding. Assistant Professor, Language Technologies Institute, CMU (2021–Present) Courtesy Appointment, Robotics Institute, CMU Research Themes : Language as a social codification of embodied experience Interpretable multimodal model training Human-robot collaboration frameworks Embodied question-answering systems Selected Trends : His recent publications show increasing focus on cross-modal attention mechanisms (Vid2Robot), error detection in toolchains (Tools Fail), and theory-of-mind reasoning in language agents (SOTOPIA). Multimodal integration spans vision, audio, and robotic control contexts (ANAVI). Labs & Collaborations : Founder of CLAW Lab (Connecting Language to Action and the World) Collaborations with Microsoft Research, Meta Inc, and CMU's REAL (Robotics, Embodied AI, Learning) community
James Glass is a Senior Research Scientist at the Massachusetts Institute of Technology (MIT) and heads the Spoken Language Systems Group within MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL). He is also affiliated with the Harvard-MIT Division of Health Sciences and Technology. His research spans automatic speech recognition, multimodal learning, and spoken language understanding, with applications in healthcare and video analysis. Education: SM and PhD in Electrical Engineering and Computer Science from MIT His work focuses on paralinguistic speech analysis, health markers in speech, and the intersection of speech and natural language processing. Recent trends emphasize audio-visual alignment, recursive reasoning, and AI applications in cognitive disorder diagnosis. Scientific awards include IEEE Fellow, ISCA Fellow, and Associate Editor for IEEE Transactions on Pattern Analysis and Machine Intelligence. His group explores unsupervised learning, speaker verification, and social text analysis. James leads the Spoken Language Systems Group at CSAIL, collaborating with institutions like IBM and Harvard-MIT Division of Health Sciences and Technology. His research integrates vision-language models, neural audio codecs, and self-supervised frameworks.
Toby Jia-Jun Li is an Assistant Professor in the Department of Computer Science and Engineering at the University of Notre Dame, where he leads the SaNDwich Lab. He also serves as the Director of the Human-Centered Responsible AI Lab in the Lucy Family Institute for Data & Society and is a Faculty Fellow at the Institute for Educational Initiatives (IEI). Previously, he was affiliated with Carnegie Mellon University's Human-Computer Interaction Institute (HCII) and GroupLens Research. Dr. Li's research spans the intersection of Human-Computer Interaction (HCI), End-User Software Engineering, Machine Learning (ML), and Natural Language Processing (NLP), with recent work focusing on addressing societal challenges in the future of work through human-AI collaborative approaches. His work has resulted in over 40 publications at premier venues including CHI, UIST, CSCW, ACL, and ICSE, with 8 papers winning Best Paper or Honorable Mention awards. His recent publications demonstrate a strong focus on human-AI collaboration across various domains, including code understanding, privacy, accessibility, and creative tools. The work shows a trajectory toward increasingly sophisticated integration of human-centered design with AI capabilities, particularly using large language models to enhance human productivity and address societal challenges. Google Research Scholar Award recipient Recipient of Yahoo! Fellowship ($100,000/year) Best Paper Award at UIST 2020 Best Paper Honorable Mention Award at CHI 2021 Best Paper Award at CSCW 2024 Best Paper Award at CHI 2025 Dr. Li actively mentors Ph.D. students and has established collaborations with Google, Microsoft Research, IBM Research, Adobe, Verizon, and J.P. Morgan. His research has been supported by NSF, Google Research Scholar Program, AnalytiXIN Initiative, Yahoo! InMind project, and J.P. Morgan. He is currently recruiting Ph.D. students and undergraduate researchers for his SaNDwich Lab, which focuses on developing interactive systems to empower individuals to create, configure, and extend AI-powered computing systems.
Rita Cucchiara is a Full Professor at the Department of Engineering 'Enzo Ferrari' of the University of Modena and Reggio Emilia. She leads the AImageLab research laboratory, part of the Artificial Intelligence Research and Innovation Center (AIRI) in Modena. Her research focuses on Computer Vision, Pattern Recognition, Machine Learning, and Multimedia, with applications in video surveillance, medical imaging, human-centered AI, and generative models. She is actively involved in interdisciplinary projects like ELIAS (European Lighthouse for AI Sustainability) and ELSA (European Lighthouse on Secure AI). Her recent roles include being elected Rector of the University of Modena and Reggio Emilia in 2025. She has organized and participated in major AI events, including workshops at NeurIPS, CVPR, and ECCV, and has contributed to advancements in multimodal models, deepfake detection, and trustworthy AI. Education details are not explicitly provided, but her extensive academic and research experience at the University of Modena underscores her expertise. She collaborates with institutions like NVIDIA, CINECA, and industry partners such as Digital Design and NVIDIA's AI Technology Center. AImageLab's projects include developing systems for medical imaging, ethical AI, and generative adversarial networks (GANs) for design surfaces. She co-organizes initiatives like the ELLIS Summer School on Large-Scale AI and contributes to policy discussions on AI ethics and societal impact. Her work spans from foundational research (e.g., vision transformers, continual learning) to applied projects (e.g., DDGan system for surface printing). Key grants and collaborations include the FAIR project and PNRR-M4C2 initiatives. She advises students and researchers in AI, with 7 PhD positions funded under national programs. Her leadership roles in AIRI and AImageLab highlight her commitment to bridging academia and industry, fostering innovation in AI-driven solutions for sustainability and healthcare.
Ira Kemelmacher-Shlizerman is a Full Professor of Computer Science at the Paul G. Allen School of Computer Science & Engineering at the University of Washington and Director of the UW Reality Lab. She also serves as a Principal Scientist at Google, where she leads the Shopping Gen AI visuals teams focusing on Virtual Try-On, 3D, and product videos. Her research spans computer vision, computer graphics, and Generative AI, with particular contributions to virtual try-on technology, 3D modeling, and augmented reality applications. Professor Kemelmacher-Shlizerman's research interests focus on Generative AI applications in visual computing. Her work bridges the gap between theoretical computer vision and practical applications, particularly in e-commerce and virtual reality. She has made significant contributions to virtual try-on technology, 3D editing with generative models, and AI applications for shopping experiences. Her research combines deep learning with traditional computer vision techniques to solve challenging problems in image and video synthesis. Her recent publications demonstrate a strong trend toward Generative AI applications for visual shopping experiences, virtual try-on technology, and 3D content creation. The work spans multiple top conferences including CVPR, SIGGRAPH, and ICCV, with a focus on practical applications of computer vision and graphics. Her research has evolved from foundational work in face reconstruction and aging to current applications in virtual shopping and 3D content generation. Google faculty award Madrona prize GeekWire Innovation of the Year Award Covers of CACM and SIGGRAPH Best student paper honorable mention at CVPR'21 Best demo runner up MobiSys'22 Senior member of IEEE Distinguished Member of ACM Professor Kemelmacher-Shlizerman has successfully tech-transferred multiple research projects to industry. She founded Dreambit, a startup acquired by Meta, and previously built and launched the Face Movies feature at Google. She currently leads Google's Shopping Gen AI visuals teams, focusing on 10x improvements to shopping journeys. Her UW Reality Lab serves as a hub for AR/VR research with industry partnerships. She has mentored numerous PhD students who have become researchers in both academia and industry, with several publications featuring student co-authors receiving recognition at top conferences. Professor Kemelmacher-Shlizerman leads the Graphics and Imaging Laboratory (GRAIL) and the UW Reality Lab, which focuses on augmented and virtual reality research with industry partnerships including Google. The labs work on cutting-edge projects in virtual try-on, 3D modeling, and immersive experiences, bridging academic research with real-world applications.
Michael J. Black is a Professor and Director at the Max Planck Institute for Intelligent Systems in Tübingen, Germany, where he leads the Perceiving Systems department and serves as Managing Director . He is also an Honorarprofessor at the University of Tübingen 's Faculty of Science . His career spans roles at Brown University (2000-2010), Xerox PARC, and academic-industry collaborations with Amazon and Meshcapade.