Xavier Serra is a Full Professor at the Department of Engineering at Universitat Pompeu Fabra (UPF), Barcelona. He is the founder and director of the Music Technology Group (MTG), and leads the UPF-BMAT Chair on AI and Music. He also coordinates the Master in Sound and Music Computing and serves as President of the Phonos Foundation. His research focuses on audio signal processing, sound and music computing, and computational musicology, emphasizing open science and open innovation. Education: BSc in Biology, University of Barcelona (1981) Master in Music, Florida State University (1983) PhD in Computer Music, Stanford University (1989) Research Interests: Audio Signal Processing Data-Driven and Knowledge-Driven Methodologies Music Information Retrieval Cultural Music Analysis (e.g., Carnatic/Turkish/Andalusian Music) Music Education Technology Notable Projects: CompMusic (ERC Advanced Grant, 2010-2017): Multicultural computational music analysis Open datasets: Freesound, Saraga, FSD50K Technologies: Reactable, Vocaloid, Essentia API Recent Trends in Articles: Focus on AI-driven audio processing (neural fingerprints, generative models), cross-cultural music analysis, and explainable music difficulty estimation. Awards: ERC Advanced Grant (2010) for CompMusic Project. Labs/Teams: Director of MTG, Phonos Foundation, and UPF-BMAT Chair. Active in open-source projects and international collaborations.
Dhruv Jain is an Assistant Professor in the Computer Science and Engineering Department at the University of Michigan , with affiliations in the School of Information and Michigan Medicine . He leads the Soundability Lab , focusing on transforming hearing into a programmable interface through human-centered AI systems. PhD from University of Washington MS from MIT Media Lab Formerly worked at Microsoft Research, Google, and Apple Research Interests intersect Human-Computer Interaction (HCI) , Audio AI , Accessibility , and Hearing Health . His work develops systems like SoundWatch (smartwatch-based sound awareness) and HomeSound (IoT sound visualization) to expand auditory access for DHH individuals and other domains. Recent Publications (2023–2025) span top venues like CHI , ASSETS , and ICMI , with trends including generative AI for real-time captioning , adaptive soundscapes , and clinical communication tools . SIGCHI Outstanding Dissertation Award (2023) William Chan Memorial Dissertation Award (2023) Best Poster Award at ASSETS 2024 Google Academic Research Award (2024) NIH Grant ($450k, 2023) William Demant Foundation Grant ($740k, 2025) Teaching includes Accessible Computing (undergraduate) and Advanced Accessibility (graduate), alongside global DIY workshops in five countries. The Soundability Lab collaborates with neuroscientists, clinicians, and Deaf scholars to create deployable systems like clinical communication tools now at Michigan Medicine and features integrated into Apple devices.
Habib Ullah is an Associate Professor in Data Science at the Norwegian University of Life Sciences (NMBU), Norway, where he conducts research at the intersection of computer vision and machine learning. He is affiliated with the Institute of Data Science under the Faculty of Science and Technology. He has previously held academic positions at COMSATS University Islamabad, Pakistan, and the University of Ha'il, Saudi Arabia, and served as a postdoctoral researcher at The Arctic University of Norway. Educational Background: PhD in Information and Communication Technology (Computer Vision), University of Trento, Italy (2011–2015) MSc in Electronics and Computer Engineering, Hanyang University, South Korea (2007–2009) BSc in Computer Systems Engineering, NWFP University of Engineering and Technology, Pakistan (2002–2006) Habib Ullah's research is primarily focused on computer vision and machine learning, with applications in aquaculture, agriculture, and human behavior analysis. He investigates underwater fish feeding sounds using audio classification, develops zero-shot learning models for recognizing unseen classes, and applies deep learning to detect stress in salmon via skin dot patterns. He also explores AI-driven controlled environment agriculture, leveraging sensors and automation for optimal crop growth. His work emphasizes practical AI solutions for real-world challenges in environmental and biological domains. The recent publications highlight a strong trend in leveraging deep learning for zero-shot and semi-supervised learning, particularly in computer vision tasks such as sea ice classification, crowd anomaly detection, and agricultural monitoring. His research spans remote sensing, biomedical signal processing, and human activity recognition, demonstrating interdisciplinary versatility. The keywords reflect a focus on robust feature representation, knowledge transfer, and model generalization. Scientific Awards and Funding: Industrial PhD grant 'Advancing Controlled Environment Agriculture AI' from The Research Council of Norway (Project number 354125, 2 million NOK, 2024) Team member (Coordinator-Participant) in the Battery Cell Assembly Twin (BatCAT) project funded by Horizon Europe (7 mEuro, 2023–2027) Development of an AI-Based Image Analysis System for Monitoring Plant Status (Funding: 1.8 mNOK, starting 2025) Habib Ullah actively supervises PhD projects and contributes to academic service through editorial and organizational roles. He has served as an Associate Editor for IEEE Access, Guest Editor for MDPI Remote Sensing, and Editor of the Springer book Machine Learning Techniques and Sensor Applications for Human Emotion, Activity Recognition, and Support (ML-SHEARS) . He has also been a Track Chair and Program Committee Member for several international conferences, reflecting his leadership in the academic community. His research is supported by significant grants and collaborative projects, indicating strong institutional and international engagement. He is involved in multiple research teams and projects, including the BatCAT project on battery manufacturing and AI applications in controlled environment agriculture with RIFT LABS AS. His lab work integrates deep learning, sensor fusion, and data analytics for environmental and biological monitoring systems.
Eric Smialek is a Senior Research Fellow at the University of Huddersfield, specializing in extreme metal vocals, popular music studies, and Taylor Swift analysis. He holds a Marie Skłodowska-Curie Actions Postdoctoral Fellowship and is affiliated with the Centre for Research in Music and its Technologies. His research explores vocal expression, cultural meaning in metal, and intersections of music with class, disability, and LGBTQ+ advocacy. Education: PhD in Genre and Expression in Extreme Metal Music (McGill University, 2016) Master's in Rethinking Metal Aesthetics (McGill University, 2009) Bachelor's from University of British Columbia (2006) Research Interests: Smialek’s work bridges music analysis and sociology, focusing on extreme metal’s vocal techniques, global social-class divides in metal scenes, and Taylor Swift’s cultural impact. He co-founded the European Taylor Swift Research Network and organized the 2024 Taylor Swift unconference in Amsterdam. Awards: Marie Skłodowska-Curie Actions Postdoctoral Fellowship Advising & Collaboration: Accepts PhD students and collaborates on projects like Extreme Metal Vocals: Musical Expression, Technique, and Cultural Meaning (2022–2025). His work has been cited in venues like Metal Music Studies and Contemporary Music Review . Labs/Teams: Member of the Centre for Research in Music and its Technologies, contributing to interdisciplinary studies in music acoustics and cultural analysis.
Richard M. Stern is a Professor of Electrical and Computer Engineering at Carnegie Mellon University (CMU), holding courtesy appointments in the Language Technologies Institute and Department of Computer Science, and serving as an Artist Lecturer in the School of Music since 2007. His interdisciplinary work bridges engineering and music technology through the School of Music's programs. Education: Ph.D. in Electrical Engineering from Massachusetts Institute of Technology (MIT), 1976 Professor Stern's research spans sound, speech, hearing, and music, with core emphases on robust speech processing in variable acoustic environments, music information retrieval, automated accompaniment, and foundational contributions to binaural perception theory. His work integrates psychoacoustic principles with machine learning to address challenges in speech recognition and human-robot interaction. Recent publications (2022-2025) reveal intensified focus on deep learning for speech enhancement in reverberant/noisy conditions, human-robot interaction scenarios, and music tagging—highlighting innovations in beamforming, source separation, and temporal modulation modeling. Awards and Honors: Fellow of the IEEE Fellow of the Acoustical Society of America Fellow of the International Speech Communication Association (ISCA) ISCA Distinguished Lecturer Allen Newell Award for Research Excellence (1992) Lutron Award for Teaching Excellence (2018) Professor Stern has advised numerous graduate students in speech and audio research, though specific names are unlisted in source materials. His grant portfolio includes significant National Science Foundation and industry-funded projects in speech technology, with leadership roles in initiatives like Interspeech 2006. He actively collaborates with CMU's Language Technologies Institute and Music and Technology program. He maintains strong ties to CMU's interdisciplinary ecosystem through the Language Technologies Institute and School of Music's Music and Technology program, contributing to research that merges acoustic engineering with musical applications.
Prof. Bernhard U. Seeber is an Extraordinary Professor at the Technical University of Munich (TUM), leading the Chair of Audio Signal Processing within the TUM School of Computation, Information and Technology. His work bridges auditory neuroscience and engineering, focusing on improving hearing aids, cochlear implants, and virtual acoustic systems. He holds affiliations with the Bernstein Center for Computational Neuroscience, Munich Institute of Biomedical Engineering, and others. Education: Studied and earned his PhD (2003) in Electrical Engineering and Information Technology at TUM. Postdoctoral research included time at UC Berkeley and the MRC Institute of Hearing Research (UK), where he pioneered studies on binaural hearing and cochlear implant optimization. Research Interests: Combines experimental and theoretical approaches to explore auditory scene analysis, binaural unmasking, and spatial hearing. Key areas include signal coding for cochlear implants, virtual acoustics, and non-destructive acoustic monitoring. His work emphasizes interdisciplinary collaboration with industry and academia. Awards: Lothar Cremer Award (2010), Emmy Noether Fellowship (2007), and recognition from the German Acoustical Society. Teaching: Offers courses on audio communication, computational neuroscience, and technical acoustics. Projects: Leads initiatives like HAPPAA and Auralization, advancing sound field synthesis and hearing aid algorithms. Current Roles: Head of Chair of Audio Signal Processing, Board Member of DEGA, and spokesperson for the ITG Technical Committee on Hearing Acoustics.
Dr. Gan Sheuo Hui serves as Lecturer in Animation at LASALLE College of the Arts, Singapore, with extensive international engagement through visiting positions at National University of Singapore (NUS), Kyoto Seika University, and the Sainsbury Institute for the Study of Japanese Arts and Cultures. Her academic profile bridges Japanese animation scholarship with Southeast Asian cultural contexts through fieldwork-driven research. Her educational qualifications include: PhD in Human and Environmental Studies from Kyoto University, Japan Master of Communication (Screen Studies) from University of Science Malaysia Bachelor of Communication (Film Studies) from University of Science Malaysia Dr. Gan's research examines Japanese anime, manga, and popular culture through lenses of authorship, censorship, and cultural hybridity, with particular focus on transnational adaptations in Southeast Asia. She integrates theoretical frameworks with empirical fieldwork involving creators, curators, and fans across Asia and the West, emphasizing practical production aspects alongside critical analysis. Her interdisciplinary approach connects film studies, media theory, and cultural anthropology. Her scholarly trajectory reveals consistent exploration of animation aesthetics (especially limited/selective animation), historical evolution of Japanese animation, and hybrid cultural forms in Southeast Asia. Key thematic threads include re-evaluating animation techniques beyond Western paradigms, analyzing global anime circulation, and documenting local adaptations of manga in Malaysia. This body of work demonstrates strong methodological integration of archival research, industry analysis, and ethnographic observation. Scientific recognition includes: Two-year Japan Society for the Promotion of Science (JSPS) Grant for postdoctoral research on Japanese anime history Sainsbury Institute for the Study of Japanese Arts and Cultures Fellowship Dr. Gan has secured competitive research funding including the JSPS grant, and maintains active international collaboration through invited lectures at institutions like Sorbonne, Tokyo University, and Keio University across 20+ countries. Her teaching innovations include NUS modules incorporating Japan-based field studies, reflecting her commitment to experiential learning in Japanese visual culture. While specific student advising isn't detailed, her pedagogy emphasizes creator-fan industry ecosystems. Her research network spans the Archive Center for Anime Studies (Niigata University), Kyoto Seika University's Manga Department, and global symposia, facilitating cross-institutional knowledge exchange on animation history and contemporary media practices.
Eduardo Mercado III is a Professor in the Department of Psychology at the University at Buffalo, College of Arts and Sciences. His research focuses on bioacoustics, cognitive psychology, and marine ecology, particularly the vocal behavior of humpback whales and its implications for understanding human impact on marine ecosystems. He is also known for his work in perceptual learning, autism spectrum disorder, and comparative cognition. Scientific Awards Guggenheim Fellowship Harvard Radcliffe Institute Fellowship Research Trends His recent publications emphasize bioacoustic analysis of humpback whale songs, including their spectral entropy, cyclical variations, and adaptive adjustments to anthropogenic noise. Additional work explores perceptual learning mechanisms in autism, neural network modeling for acoustic classification, and cognitive processes in canines and rodents. Projects Mercado’s “Singers as Sentinels” project combines acoustic analysis of humpback whale songs with public awareness initiatives about ocean noise pollution. The project will produce a book, Why Whales Sing and Dolphins Don’t , and a web-based interface for public engagement.
Professor Marie Roch is a distinguished faculty member in the Department of Computer Science at San Diego State University within the College of Sciences . Her groundbreaking research bridges Bioacoustics and Machine Learning , focusing on advanced algorithms for automated detection, classification, and analysis of marine mammal vocalizations using passive acoustic monitoring. Core research in marine bioacoustic signal processing and deep learning applications for echolocation click detection Published extensively in Journal of the Acoustical Society of America , Biological Reviews , and IEEE Transactions Developed deep learning frameworks for whale whistle extraction without human annotation Created open-source tools like Silbido Profundo for automated marine mammal call analysis Marie's work has been supported by over $3 million in grants from the DOD Office of Naval Research , Bureau of Ocean Energy Management , and Human Frontier Science Program . She actively mentors graduate students and serves on numerous thesis committees, with recent advisees working on deep learning for baleen whale calls and terrestrial animal recognition . Her Marine Acoustic Research Lab (MAR Lab) leads in developing the Tethys metadata workbench for ocean acoustic data management.
Dr. Armin Mustafa is an Associate Professor in Computer Vision and AI at the University of Surrey, where he holds a prestigious Royal Academy of Engineering Research Fellow position. He is affiliated with the Centre for Vision, Speech and Signal Processing (CVSSP), the School of Computer Science and Electronic Engineering, and the Surrey Institute for People-Centred Artificial Intelligence (PAI). His research focuses on developing AI systems for visual understanding of complex dynamic scenes, with applications in entertainment, autonomous systems, and augmented/virtual reality. Dr. Mustafa completed his PhD in general dynamic scene reconstruction from multi-view videos in 2016 from the University of Surrey under the supervision of Prof. Adrian Hilton. Prior to his doctoral studies, he worked for three years (2010-2013) at Samsung Research Institute in Bangalore, India, in the field of Computer Vision. His research expertise spans Computer Vision, Scene Understanding, 3D/4D Vision, Virtual Reality, Light Fields, Machine Learning, Video Captioning, Augmented Reality, Artificial Intelligence, and Audio-visual Video Understanding. Dr. Mustafa has pioneered advances in 4D vision, NLP, and Scene Understanding over the past decade, with a particular focus on enabling machines to model and interpret real-world environments for socially beneficial applications. His work bridges theoretical advances in computer vision with practical applications in media production, virtual reality, and autonomous systems. Analysis of Dr. Mustafa's recent publications reveals a strong focus on multimodal learning, particularly the integration of audio and visual information for scene understanding. His work spans diverse areas including shadow detection and removal, audio event classification, video captioning, person image generation, and dynamic scene reconstruction. A notable trend is his exploration of transformer architectures for both vision and audio tasks, as well as the application of self-supervised learning techniques to reduce dependency on labeled data. Dr. Mustafa has received numerous prestigious awards: 2018 - Research Fellowship, The Royal Academy of Engineering, UK 2017 - Young Researcher award, CVPR 2016 - Doctoral Consortium grant, CVPR 2015 - BMVA travel grant for ICCV 2014 - Set-Squared Research to Innovator grant 2013 - Overseas Research Scholarship, FEPS, The University of Surrey 2010 - Cadence Silver Medal, Indian Institute of Technology, Kanpur As a dedicated mentor, Dr. Mustafa supervises several PhD students working on cutting-edge topics including multi-person reconstruction, audio-visual scene understanding, and automatic storyboard generation. His research is supported by significant grants including a £15 million UKRI Prosperity Partnership with the BBC (AI4ME), a 5-year Royal Academy of Engineering fellowship (4D Vision for Perceptive Machines), and multiple projects with industry partners such as Figment Productions and Foundry. Dr. Mustafa is an active member of the Centre for Vision, Speech and Signal Processing (CVSSP), one of the world's leading research centers in vision, speech, and signal processing. He also contributes to the Surrey Institute for People-Centred Artificial Intelligence (PAI), where he serves as a Surrey AI Fellow. His work often involves collaboration with industry partners and other academic institutions across Europe.
Bobby Lee Townsend Sturm JR is an Associate Professor at KTH Royal Institute of Technology, leading the MUSAiC project (ERC-2019-COG). He holds a PhD in Electrical and Computer Engineering from UC Santa Barbara (2009), followed by postdoctoral research at LAM, Paris 6, and academic roles at Aalborg University and Queen Mary University of London. His research focuses on AI ethics in music, generative AI for music, and folk music preservation. Current roles at KTH include teaching and supervising in Machine Learning, Music Informatics, and AI Ethics. He has pioneered AI music generation challenges (e.g., 2020 Double Jigs Challenge) and investigates societal impacts of AI on traditional music cultures. His work bridges technical innovation with cultural and ethical considerations, addressing issues like data colonialism, algorithmic bias, and human-AI collaboration in creative contexts. Education: PhD (UCSB, 2009), Postdoc (Paris 6), Academic appointments at Aalborg University (2010–2014) and Queen Mary University (2014–2018) Key Projects: MUSAiC (ERC), Virtual Session System for Irish Music, Traditional Music Dataset Analysis Teaching: Courses in Machine Learning, Music Acoustics, and ICT Innovation Publications span peer-reviewed journals and conferences, emphasizing ethical AI, music generation, and interdisciplinary research in MIR (Music Information Retrieval). He actively collaborates with musicians, anthropologists, and technologists to ensure culturally informed AI development.
Professor Guy Brown is Chair of Computer Science at the University of Sheffield's School of Computer Science. He holds a BSc in Applied Science (1984), PhD in Computer Science (1992), and MEd in Teaching and Learning (1997). His research focuses on Computational Auditory Scene Analysis (CASA), noise-robust speech recognition, auditory modeling, and binaural processing. Research interests include: Machine hearing systems for sound source separation Reverberation-robust speech processing Auditory scene analysis models for normal/impaired hearing Applications in robotics and healthcare technologies Publication trends show recent focus on deep learning approaches for biomedical applications including sleep apnea detection, respiratory sound analysis, and multimodal health monitoring systems using neural networks. Honors include: University Senate Award for Excellence in Teaching (2014) Microsoft Software Engineering Innovation Award (2013) He leads doctoral supervision for 15+ students and has secured research funding from EPSRC, Innovate UK, EU FP7, and AHRC. Manages the Speech and Hearing research group and has held visiting positions at international institutions including LIMSI-CNRS and ATR Japan.
Chenliang Xu is an Associate Professor in the Department of Computer Science at the University of Rochester, affiliated with the Goergen Institute for Data Science and Artificial Intelligence (GIDS-AI). His research focuses on computer vision, audio-visual learning, and trustworthy AI. He holds a PhD from the University of Michigan (2016), with prior degrees from Nanjing University of Aeronautics and Astronautics and the University at Buffalo. Notable awards include the Best Paper Award at ACCV 2024 and the James P. Wilmot Distinguished Professorship. His work spans interdisciplinary topics such as video understanding, multimodal reasoning, and robust AI. Key research contributions include audio-visual scene synthesis, bias mitigation in models, and applications in public health. He has secured over $3M in grants, including NIH funding for AI-driven video description tools and public health initiatives. Prof. Xu advises a dynamic research group with 11 PhD students and numerous collaborators. His lab explores cutting-edge projects like egocentric audio-visual understanding, generative AI for avatars, and multimodal defense mechanisms. He teaches courses in machine vision, deep learning, and advanced computer vision.
Mark Plumbley is a Professor of Signal Processing at the Centre for Vision, Speech and Signal Processing (CVSSP) within the School of Computer Science and Electronic Engineering at the University of Surrey. He holds an EPSRC Fellowship in 'AI for Sound' and has led major research initiatives, including the DCASE challenges. His work focuses on AI-driven analysis of acoustic scenes and events, with contributions to machine learning, audio source separation, and sparse representations. Previously, he was Director of the Centre for Digital Music at Queen Mary University of London and Head of the School of Computer Science at Surrey. Education: PhD in Neural Networks (1991). Academic roles include Professorships at King’s College London (1991–2002) and Queen Mary University of London (2002–2014). Research spans audio event detection, sound scene classification, and generative AI for audio synthesis. He leads projects like the EPSRC-funded 'Making Sense of Sounds' and 'Musical Audio Repurposing using Source Separation', and co-edited the Springer book on Computational Analysis of Sound Scenes and Events. Research Interests: AI for Sound: Machine learning applied to real-world audio analysis. Acoustic Scene and Event Recognition: Developing models for sound classification and localization. Generative Audio Models: Text-to-audio systems and diffusion models for sound synthesis. Healthcare Applications: Audio-based diagnostics and bioacoustic signal processing. Grants and Awards: EPSRC Fellowships, EU-funded networks (SpaRTaN, MacSeNet), and Fellowships from IET and IEEE. Notable awards include the IEEE Young Author Best Paper Award (co-authored with students) and leadership in the DCASE community. Labs and Collaborations: CVSSP at Surrey, collaborations with BBC R&D, and interdisciplinary projects on urban soundscapes and noise pollution (UK Acoustics Network Plus).
Slim Essid is a Full Professor at Télécom Paris, leading the Audio Data Analysis and Signal Processing (ADASP) group. He holds a Doctorat (Ph.D.) and Habilitation from Université Pierre et Marie Curie (UPMC). With 15+ years of research experience, he has advised 15 PhD graduates and currently co-advises 10 others. His work focuses on machine learning, signal processing, and multimodal systems, publishing over 150 peer-reviewed papers. He serves as a reviewer for top journals/conferences (e.g., IEEE Transactions) and research funding agencies. Education: State Engineering Degree, École Nationale d’Ingénieurs de Tunis (2001) M.Sc. (D.E.A.) in Digital Communication Systems, École Nationale Supérieure des Télécommunications, Paris (2002) Ph.D., Université Pierre et Marie Curie (2005) Habilitation (HDR), UPMC (2015) Research Interests: Multimodal learning, self-supervised representations, audio-visual segmentation, music structure analysis, domain generalization, and speech enhancement. Recent publications highlight innovations like TACO (training-free sound-prompted segmentation) and CLOUDS (domain-generalized semantic segmentation framework using foundation models). His work bridges audio processing with vision and language models, emphasizing unsupervised/zero-shot approaches. Key achievements include state-of-the-art methods in sound event detection, speaker diarization, and music segmentation. He collaborates with 14 post-docs and leads projects funded by French/EU agencies.