Chris Donahue is an Assistant Professor in the Computer Science Department at Carnegie Mellon University . He also serves as a part-time Research Scientist at Google DeepMind on the Magenta team. His work focuses on leveraging generative AI to enhance human creativity, particularly in music. Education: PhD in Computer Science (UC San Diego), Postdoctoral Scholar (Stanford University) His research spans controllable generative modeling of music and audio , with a focus on real-time interactive systems. Projects like Piano Genie , Beat Sage , and Copilot Arena demonstrate his commitment to real-world deployment. His Generative Creativity Lab (G-CLef) explores AI applications beyond music, including programming and natural language. Recent publications highlight advancements in multimodal music evaluation , real-time adaptation , and AI-driven sound morphing . He co-developed Magenta RealTime , an open-weight real-time music generation model, and MusicFX DJ Mode . Scientific Awards: Best Paper Award (top 1) at NAACL Student Research Workshop 2025 Best Paper Award (top 1% of submissions) at CHI 2025 Best Paper Runner-up at ISMIR 2021 He co-advises PhD students like Wayne Chi (NDSEG Fellow) and mentors Irmak Bukey . His lab receives support from the AIxArts incubator fund at CMU .
Professor Bernd Möbius is a leading academic in Phonetics and Phonology at the Department of Language Science and Technology, Saarland University. His research bridges phonetic theory with speech technology applications, focusing on text-to-speech systems, prosody modeling, and computational simulations of speech processes. Current research projects: DFG SFB 1102, C1: Information density and phonetic structure predictability DFG SFB 1102, C4: Slavic intercomprehension and surprisal theory (INCOMSLAV) Research Themes: Key areas include text-to-speech synthesis, speech prosody analysis, experimental methods in speech production/perception, information density in phonetics, and cross-linguistic studies of Slavic-Germanic languages. Scientific Contributions: Recent work explores Parkinson-induced dysarthria detection, breath noise acoustics, surprisal-driven speech behaviors, multilingual BERT models for idiomaticity, and perceptual consequences of acoustic adjustments.
WANG Ye is an Associate Professor in the Department of Computer Science at the School of Computing, National University of Singapore (NUS). He holds a PhD in Information Technology from Tampere University of Technology, Finland, and has been a tenured faculty member at NUS since 2002, following his industry research role at Nokia Research Center. He is the director of the Sound and Music Computing Lab at NUS, leading cutting-edge research in AI-driven music and health technologies. PhD, Information Technology, Tampere University of Technology, Finland (2002) MSc, Telecommunications, Braunschweig University of Technology, Germany (1993) BSc, Telecommunications, South China University of Technology, China (1983) His research is centered on Sound and Music Computing for Human Health and Potential (SMC4HHP) , with a focus on eHealth, eLearning, mobile/wearable computing, and music information retrieval. His work spans AI for stroke rehabilitation, language learning through singing, singing voice synthesis, and automatic music transcription. He has pioneered systems like SLIONS (language learning via karaoke), CocoLyricist (AI co-creation for stroke recovery), and SinTechSVS (expressive singing voice synthesis). The latest articles highlight a strong trend in AI-driven music and health technologies , particularly in controllable lyric generation, singing voice synthesis, automatic pronunciation assessment, and multimodal music transcription. The research increasingly integrates large language models, explainable AI, fairness, and real-world deployment, reflecting a shift from theoretical exploration to practical, human-centered applications in healthcare and education. Dr. Wang has received numerous scientific honors, including: Best Paper Awards at ACM MM, ISMIR, IEEE ISM, and CHI First Prize, Asia Pacific Assistive, Rehabilitative, and Therapeutic Technologies Challenge (2015) Faculty Teaching Excellence Award, NUS School of Computing (2024) Top Paper Award, ACM Multimedia 2022 AI in Medicine Collaborative Grant for CocoLyricist project He has supervised over 11 PhD and 20 MComp students and is currently guiding six PhD candidates. His grants come from MOE, NRF, A*STAR, Nokia, and Smule. He has served as General Chair of ISMIR2017 and TPC Co-Chair of ICOT2017, and is on the editorial boards of IEEE Transactions on Multimedia and Journal of New Music Research. He has also developed and taught the first course on Sound and Music Computing in Singapore. Dr. Wang leads the Sound and Music Computing Lab (SMC Lab) , a multidisciplinary team exploring the synergy of music computing, AI, mobile technology, and cloud systems for health and education. The lab actively collaborates with medical institutions such as NUS Yong Loo Lin School of Medicine, Singapore General Hospital, and Harvard Medical School, and is currently working on projects in AI-supported language learning, stroke rehabilitation, and intelligent music interfaces.
Catherine Lai is a Reader (~Associate Professor) in the Department of Linguistics and English Language at the University of Edinburgh, with strong affiliations to the Centre for Speech Technology Research (CSTR) and the Institute for Language, Cognition and Computation (ILCC) in the School of Informatics. She is based in the School of Philosophy, Psychology and Language Sciences and is actively involved in research, teaching, and academic service. Department: Department of Linguistics and English Language School: School of Philosophy, Psychology and Language Sciences Research Institutes: Centre for Speech Technology Research, Institute for Language, Cognition and Computation Email: C.Lai@ed.ac.uk Her research centers on the role of prosody—non-lexical aspects of speech—in spoken communication. She investigates how prosody contributes to discourse structure, information structure, and affect in dialogue, using interdisciplinary methods from linguistics and machine learning. Her work bridges theoretical linguistics and practical speech technology, aiming to improve spoken language understanding and synthesis systems. She is particularly interested in how prosody shapes listener expectations and how affect and topic are expressed and perceived in conversation. Her recent publications reflect a strong focus on self-supervised learning in speech models, emotion recognition, ASR error correction using large language models, cognitive state classification, and ethical considerations in language technology. She explores topics such as the uncanny valley in synthetic speech, gender expression through voice, and community-centered development of language technologies. Prize from Scopus Profile Catherine Lai has supervised several PhD students, including Leimin Tian and Yuanchao Li, and has been involved in significant research projects, such as a Toyota-funded initiative on spoken dialogue for robot companions. She has secured multiple grants and leads a research agenda that integrates theoretical inquiry with real-world applications in assistive technologies and social science. Her academic service includes organizing major conferences like Interspeech and UK and Ireland Speech. She is a key member of research teams at CSTR and ILCC, collaborating across disciplines to advance the understanding of spoken communication and the development of robust, ethical speech technologies.
Jens Edlund is a Professor at KTH Royal Institute of Technology's Division of Speech, Music and Hearing. His research focuses on speech technology, dialogue systems, prosody, and evolutionary phonetics. He has contributed to foundational work on speech synthesis, conversational interaction, and multimodal corpora like the D64 corpus. Key projects include the MonAMI Reminder system and analysis of primate vocalizations to understand speech evolution. Edlund has collaborated extensively with global researchers, producing over 150 peer-reviewed works. His work integrates computational methods with linguistic and biological insights, emphasizing human-like dialogue systems and cross-species vocal analysis. Education: Ph.D. in Speech Technology (2011, KTH) Grants: Multiple EU and Swedish Research Council grants for speech technology and interdisciplinary studies Research labs include the KTH Speech, Music and Hearing Lab and collaborations with institutions like Max Planck Institute for Evolutionary Anthropology. Current work explores evolutionary origins of speech biomechanics and AI-driven speech synthesis evaluation.
University of Illinois Urbana-ChampaignUnited States
Pasquale Bottalico serves as Associate Professor in the Department of Speech and Hearing Science at the University of Illinois, with dual appointments as Associate Professor at the Center for Latin American and Caribbean Studies and Affiliate Faculty in the School of Music. His unique interdisciplinary profile bridges engineering, music performance, and speech science, reflecting his dual academic training and professional artistry. His educational foundation includes: Bachelor's in Telecommunications Engineering from Univeristà Mediterranea di Reggio Calabria, Italy Concurrent Opera Singing degree from F. Cilea Music Academy, Reggio Calabria Master's in Telecommunications Engineering from Politecnico di Torino, Italy Ph.D. in Metrology specializing in acoustics measurement uncertainty and classroom acoustics Dr. Bottalico's research centers on vocal load quantification and professional voice techniques , with significant contributions to understanding vocal fatigue in teachers and singers. His work spans Speech Intelligibility in educational environments, Room Acoustics for performance and learning spaces, and Musical Acoustics of historical vocal styles. A distinctive thread throughout his research examines how acoustic conditions modulate voice production and perception, increasingly incorporating virtual reality and bone conduction technologies for innovative assessment and intervention approaches. His Colombian vocal health study demonstrates cross-cultural applications of his work. Analysis of his 2023-2025 publications reveals three dominant research trajectories: (1) The impact of noise and dysphonia on children's speech processing in educational settings, using multimodal assessment including EEG; (2) Virtual reality applications for voice production research and therapeutic intervention; (3) Cross-cultural validation of vocal fatigue metrics and development of biofeedback systems. His work consistently bridges engineering precision with clinical applicability, particularly for professional voice users in challenging acoustic environments. No scientific awards were documented in the available information. While specific advising relationships aren't detailed, his research collaborations span international institutions including Colombian and Italian universities, suggesting graduate mentorship in interdisciplinary projects. No grant information was provided, though his systematic reviews and cross-cultural studies imply externally funded research activities. Though no dedicated laboratory is specified, his virtual reality voice studies and acoustic parameter assessments suggest affiliations with audio engineering facilities and voice clinics, likely through the Speech and Hearing Science department's research infrastructure.
Xavier Serra is a Full Professor at the Department of Engineering at Universitat Pompeu Fabra (UPF), Barcelona. He is the founder and director of the Music Technology Group (MTG), and leads the UPF-BMAT Chair on AI and Music. He also coordinates the Master in Sound and Music Computing and serves as President of the Phonos Foundation. His research focuses on audio signal processing, sound and music computing, and computational musicology, emphasizing open science and open innovation. Education: BSc in Biology, University of Barcelona (1981) Master in Music, Florida State University (1983) PhD in Computer Music, Stanford University (1989) Research Interests: Audio Signal Processing Data-Driven and Knowledge-Driven Methodologies Music Information Retrieval Cultural Music Analysis (e.g., Carnatic/Turkish/Andalusian Music) Music Education Technology Notable Projects: CompMusic (ERC Advanced Grant, 2010-2017): Multicultural computational music analysis Open datasets: Freesound, Saraga, FSD50K Technologies: Reactable, Vocaloid, Essentia API Recent Trends in Articles: Focus on AI-driven audio processing (neural fingerprints, generative models), cross-cultural music analysis, and explainable music difficulty estimation. Awards: ERC Advanced Grant (2010) for CompMusic Project. Labs/Teams: Director of MTG, Phonos Foundation, and UPF-BMAT Chair. Active in open-source projects and international collaborations.
Andrew Head is an Assistant Professor at the University of Pennsylvania in the Computer and Information Science department. His research focuses on human-computer interaction, programming, and reading, particularly in developing technology for interactive reading and reasoning. He advises PhD students Alyssa Hwang, Hita Kambhamettu, Litao Yan, Jeffrey Tao, and Jessica Shi, and co-leads the Penn Human-Computer Interaction (Penn HCI) group with Danaé Metaxa. His work is published in top venues like ACM CHI, UIST, and ICSE. Research Interests: Andrew's work bridges interactive systems with programming environments, aiming to enhance how scientists and programmers interact with their tools. Key areas include AI-assisted code understanding, math notation accessibility, and medical note interpretation through interactivity. He employs user studies to identify needs and builds interactive systems to address them. Recent Article Trends: His publications emphasize systems-centric HCI approaches, integrating AI into code and document interfaces. Topics span code explanation (e.g., Ivie), property-based testing (e.g., Tyche), math notation augmentation (e.g., FreeForm), and medical informatics (e.g., Explainable Notes). Recent work also explores notebook environments (e.g., Bolt-on, Tyche) and literate programming (e.g., Colaroid). Scientific Awards: Distinguished Paper Award, ICSE 2024 Best Paper Awards at CHI (2024, 2023, 2022, 2019), UIST (2023, 2018), and others. Nominated for Best Paper at CHI 2023 and VL/HCC 2015. Advising & Grants: Andrew advises multiple PhD students and has secured significant grants, including a $1M NSF award for 'Property-based Testing for the People' (2024). His group collaborates with Penn’s MindCORE center and teams like PLClub and PennNLP. Labs & Teams: He co-leads the Penn HCI group, which works closely with other Penn research teams and labs. The group focuses on creating interactive tools that enhance scientific and programming workflows, supported by grants and cross-institutional partnerships.
Elena Maria Baralis is a Full Professor at the Department of Control and Computer Science (DAUIN) at the Polytechnic University of Turin. She serves as Pro-Rector, member of the Board of Directors (without voting rights), member of the Academic Senate (without voting rights), and coordinator of the University's Permanent Observatory for Monitoring the Academic Sector. She chairs the Control and Computer Engineering Department and previously chaired the Computer Engineering School from October 2012 to October 2018. Her research interests focus on database systems and data mining, specifically explainable AI, bias detection in data analytics, and machine learning algorithms for big data. Her work spans various application domains including predictive maintenance, Industry 4.0, and healthcare. Recent publications demonstrate her expertise in speech processing, bias mitigation, and innovative neural network architectures like Kolmogorov-Arnold Networks. Her research output shows a clear trend toward addressing fairness and explainability in AI systems while exploring novel approaches to speech and language understanding. Professor Baralis has received significant recognition including becoming a Fellow of the Academy of Sciences of Turin in 2017. She has served as Editor-in-Chief for IEEE Internet of Things Journal (2016-2019) and Knowledge and Information Systems (2014-present). She actively mentors doctoral students including Claudio Savelli (researching Machine Unlearning), Eleonora Poeta, Giuseppe Gallipoli, Alkis Koudounas, and others. Her research is supported by numerous projects including AI4CTI (Artificial Intelligence for Cyber Threat Intelligence, 2025-2028), Smart manufacturing driven by Machine Learning in Industry 4.0 (2019-2020), and I-REACT (2016-2019).
Dr. Edyta Chlebowska is an Assistant Professor at the Cyprian Norwid Research Center, John Paul II Catholic University of Lublin. Specializing in the visual art and literary legacy of Cyprian Norwid, she explores intersections between Polish Romanticism, art history, and cultural studies. Her work extends to album art analysis and interdisciplinary connections with music (e.g., Pink Floyd, Czesław Niemen). Her research focuses on Norwid's self-portraits, eschatological themes in his drawings, and comparative studies with 19th-century Warsaw painting. She reconstructs lost artworks, documents private collections (e.g., Konstancja Górska, Victor Gomulicki), and examines Norwid's reception in Florence and Odessa. Key projects include translating Studia Norwidiana and critical editing of Norwid's complete works. As editor-in-chief of multiple On Cyprian Norwid. Studies and Essays volumes and organizer of conferences like Niemen Non Stop and Colloquia Norwidiana , she bridges academic rigor with public engagement. Her publications span peer-reviewed books, journal articles, and popular science pieces, emphasizing Norwid's unpublished works and symbolic programs.
Jianjing Kuang is an Associate Professor of Linguistics at the University of Pennsylvania, affiliated with the School of Arts and Sciences. As Director of the Penn Phonetics Laboratory, they lead research in phonetics, laboratory phonology, and tonal language studies. Kuang holds a Ph.D. from UCLA (2013) and specializes in the interplay between production and perception in speech, particularly focusing on tonal systems, prosody, and cross-linguistic fieldwork. Their work integrates behavioral experiments, corpus studies, and computational modeling to explore phonological contrasts, voice quality, and sound change. Key affiliations include MindCORE and the Center for East Asian Studies. Research interests include multidimensional cues in tone processing, glottal articulations, mapping production-perception relationships, and prosodic sentence processing. Kuang has conducted fieldwork on languages such as Yi, Q’anjob’al, and Mandarin. Their studies address topics like tonal splitting in Yi, cue-changing in Korean stops, and prosodic patterns in Mayan languages. Recent work explores voice quality’s role in pitch perception and prosodic boundary detection. Awards and grants are not explicitly listed, but Kuang’s extensive publications (over 50 peer-reviewed papers) reflect sustained academic impact. They advise graduate students in phonetics and phonology, with active collaborations in computational linguistics and speech technology. The Penn Phonetics Laboratory, under Kuang’s direction, emphasizes interdisciplinary research bridging experimental and theoretical phonetics.
Ricardo Gutierrez-Osuna is a Professor in the Department of Computer Science and Engineering at Texas A&M University, part of the College of Engineering. He leads the PSI Lab and focuses on machine learning, speech processing, and digital health applications. His research spans topics like wearable sensors, foreign accent conversion, and physiological monitoring. Education: Ph.D. (Computer Engineering, NC State, 1998), M.S. (Computer Engineering, NC State, 1995), B.S. (Electrical Engineering, Universidad Politécnica de Madrid, 1992). Research interests include intelligent sensors, speech processing, machine learning, neuromorphic computation, and mobile robotics. His work bridges computer science and biomedical engineering, with applications in health monitoring and human-computer interaction. Awards: NSF CAREER Award (2002) Ramón y Cajal Award (2005-2010) Texas A&M Barbara and Ralph Cox Fellow (2009) Multiple teaching awards (2009-2010) His lab develops innovative technologies like stress-detecting wearables, biofeedback games, and systems for non-native speech improvement. He collaborates on projects involving voice conversion, glucose prediction algorithms, and multi-modal sensing devices.
Anil K. Jain is a University Distinguished Professor at Michigan State University, where he has taught and conducted research for over 50 years. His work focuses on Pattern Recognition , Biometrics , and Machine Learning , with foundational contributions to fingerprint, face, and palmprint recognition. B.S., Indian Institute of Technology, Kanpur (1969) M.S. and Ph.D., The Ohio State University (1970, 1973) in Electrical Engineering His research spans Computer Vision , Deep Learning , and Biometric Security , addressing challenges in adversarial robustness , demographic bias , and generative models . Recent publications emphasize transformer-based architectures , domain adaptation , and contactless biometric systems . Scientific awards include: Inductee, National Academy of Engineering (2016) Inductee, The World Academy of Sciences (2019) BBVA Foundation Frontiers of Knowledge Award (2025) Fellowships: Guggenheim, Humboldt, Fulbright Doctor Honoris Causa: 3 universities He has authored seminal works like Introduction to Biometrics and Handbook of Face Recognition , and served as Editor-in-Chief of IEEE Transactions on Pattern Analysis and Machine Intelligence . His leadership in Forensic Science includes roles on the Defense Science Board and AAAS study teams.
Kirsten Yri is an Associate Professor of Musicology at Wilfrid Laurier University, serving as Chair of Music Promotions and Appointments, and Coordinator for Music History, Theory, and Critical Analysis. She also oversees the Master of Music: Collaboration, Curation, and Creative Performance program. Her research focuses on medievalism in contemporary music and popular culture, the early music revival of the 20th century, and representations of gender in opera and music within colonial/postcolonial contexts. Yri’s research interests include analyzing how medieval themes are reinterpreted in modern contexts (e.g., film, heavy metal, and stage productions), examining the ideological underpinnings of early music revival movements, and exploring the intersection of music with gender, nationalism, and global cultural politics. Her work often critiques Western-centric narratives and highlights marginalized voices in music history. Her publications span articles in journals and edited volumes, with a focus on composers like Carl Orff and Hildegard of Bingen, as well as groups like Corvus Corax and Anonymous 4. She has contributed to interdisciplinary projects such as The Oxford Handbook of Music and Medievalism (2020). While no scientific awards are cited, her work is widely recognized in musicology and cultural studies. Yri’s teaching and research emphasize collaborative, interdisciplinary approaches to music history, blending academic rigor with creative curation. She speaks English, Swedish, and German, enhancing her ability to engage with international scholarly resources and communities.
Professor Peter F. Driessen is a faculty member in the Department of Electrical and Computer Engineering at the University of Victoria, with a cross-appointment in the School of Music. He holds a BSc and PhD from the University of Victoria and is a Professional Engineer (PEng). His research focuses on communication systems, signal processing, control, and interdisciplinary projects in computer music and wireless technologies. Key areas include audio/video signal processing, radio propagation, sound recording, and multimedia systems. He leads the University of Victoria Propagation Laboratory, which explores radio wave propagation and Amateur radio integration with engineering education. His work spans theoretical research and applied projects like ECOSat satellite systems, software-defined radio (SDR), and innovative musical instruments such as the Radio Drum. He supervises undergraduate and graduate projects in these domains through ELEC 499 courses. Notable contributions include the APEGBC Editorial Board Award for Best Paper (2002) and patents in wireless networking and signal processing. His teaching includes courses in signal analysis and electromagnetics, and he collaborates on interdisciplinary programs like the Music/Computer Science degree. Education: BSc in Electrical Engineering, University of Victoria PhD in Electrical Engineering, University of Victoria Research Interests: Audio and video signal processing for music and media Software-defined radio and Amateur radio technologies Satellite communication and ground station development Gesture-based interfaces and musical instrument design Error mitigation in streaming audio/video Optical and microwave-photonic systems Labs & Collaborations: Propagation Laboratory (radio wave research) UVic Experimental Radio Group (Amateur radio club) UVic Satellite Design Team (ECOSat projects) UVic Centre for Aerospace Research Grants & Awards: APEGBC Editorial Board Award (2002) Multiple US patents in wireless systems and signal processing