Shrikanth (Shri) Narayanan is a University Professor and holder of the Niki and Max Nikias Chair in Engineering at the University of Southern California (USC), serving as the inaugural Vice President for Presidential Initiatives. He leads the Signal Analysis and Interpretation Lab (SAIL) and holds joint appointments in Computer Science, Linguistics, Psychology, Neuroscience, Pediatrics, and Otolaryngology-Head and Neck Surgery. His research focuses on speech and audio processing, behavioral signal processing, and real-time MRI of speech production, with applications in healthcare, education, and technology. Education: B.E. in Electrical Engineering from College of Engineering, Guindy (Chennai, India, 1988); M.S., Engineer, and Ph.D. in Electrical Engineering from UCLA (1990, 1992, 1995). Research interests span computational linguistics, machine learning, and multimodal human behavior analysis. He pioneered technologies for speech biomarkers in mental health, real-time MRI of speech production, and wearable sensor systems for longitudinal health studies. His work in speech emotion recognition, forensic interviews, and clinical applications has been recognized through over 40 awards, including the IEEE Flanagan Award and ISCA Medal. He has published extensively in journals like Proceedings of the IEEE , Journal of the Acoustical Society of America , and PLOS One . Key Grants: NSF CAREER, Okawa Research, IBM Faculty, Google/Amazon awards. Labs/Teams: Signal Analysis & Interpretation Lab (SAIL), USC Information Sciences Institute (ISI), Google Visiting Faculty Researcher.
Brett Welch is an Assistant Professor and Biostatistician at the University of Houston with dual appointments in the Department of Communication Sciences and Disorders and the Department of Psychological, Health, and Learning Sciences within the College of Liberal Arts and Social Sciences. Dr. Welch's educational background includes: Ph.D. in Communication Science and Disorders from the University of Pittsburgh M.S. in Communication Sciences and Disorders from the University of Texas Rio Grande Valley B.S. in Linguistics from the University of Texas at Austin B.A. in Communication Sciences and Disorders – Speech-Language Pathology from the University of Texas at Austin Dr. Welch's research sits at the intersection of psychology and communication. He seeks to understand how psychological processes like personality and psychopathology influence communication behaviors, and conversely, how communication behaviors impact psycho-social well-being. His work leverages quantitative methods, cross-sectional studies, and intensive longitudinal designs to interrogate the dynamic interplay between psychological processes and voice and communication behaviors. This research aims to advance understanding of functional voice disorders, improve assessment of psychopathology, and promote well-being for people with and without communication disorders. Dr. Welch has received several notable awards: Honors Council of Academic Programs in Communication Sciences and Disorders PhD Scholarship (2024) New Century Scholars Doctoral Scholarship from the American Speech-Language-Hearing Foundation (2023) New Investigator Forum from The Voice Foundation (2021) As an academic, Dr. Welch serves as an ad hoc peer reviewer for various scientific journals in speech-language pathology and psychology. He is also a mentor in the Students to Empowered Professional Mentoring Program and a member of the American Speech-Language-Hearing Association. His teaching includes Special Problems: Introduction to Quantitative Modeling (COMD 8398). Dr. Welch maintains an active research program focused on the intersection of psychology and communication, with particular emphasis on voice disorders and their psychological correlates. His work combines biostatistical approaches with clinical insights to advance understanding in this specialized field.
Ricardo Gutierrez-Osuna is a Professor in the Department of Computer Science and Engineering at Texas A&M University, part of the College of Engineering. He leads the PSI Lab and focuses on machine learning, speech processing, and digital health applications. His research spans topics like wearable sensors, foreign accent conversion, and physiological monitoring. Education: Ph.D. (Computer Engineering, NC State, 1998), M.S. (Computer Engineering, NC State, 1995), B.S. (Electrical Engineering, Universidad Politécnica de Madrid, 1992). Research interests include intelligent sensors, speech processing, machine learning, neuromorphic computation, and mobile robotics. His work bridges computer science and biomedical engineering, with applications in health monitoring and human-computer interaction. Awards: NSF CAREER Award (2002) Ramón y Cajal Award (2005-2010) Texas A&M Barbara and Ralph Cox Fellow (2009) Multiple teaching awards (2009-2010) His lab develops innovative technologies like stress-detecting wearables, biofeedback games, and systems for non-native speech improvement. He collaborates on projects involving voice conversion, glucose prediction algorithms, and multi-modal sensing devices.
Max Planck Institute of Colloids and InterfacesGermany
Shrikanth (Shri) Narayanan is University Professor and Niki & C. L. Max Nikias Chair in Engineering at the University of Southern California (USC), with appointments spanning Electrical & Computer Engineering, Computer Science, Linguistics, Psychology, Neuroscience, Pediatrics, and Otolaryngology-Head & Neck Surgery. He serves as Research Director of the Information Sciences Institute and Director of the Ming Hsieh Institute. PhD in Electrical Engineering (UCLA, 1995) Engineer and MS in Electrical Engineering (UCLA, 1992 and 1990) BE in Electrical Engineering (Anna University, India, 1988) His interdisciplinary research focuses on human-centered signal processing and machine intelligence , addressing societal challenges in health, education, defense, and media arts. Key areas include: Behavioral signal processing Affective computing Multimodal signal processing Computational speech science Biomedical applications Scientific Awards : IEEE James L. Flanagan Speech and Audio Processing Award (2025) Edward J. McCluskey Technical Achievement Award (2024) ISCA Medal for Scientific Achievement (2023) Claude Shannon-Harry Nyquist Technical Achievement Award (2023) ACM ICMI Sustained Accomplishment Award (2020) USC Distinguished Faculty Service Award With over 1,000 publications and 19 patents , his work has been commercialized through startups like Behavioral Signals Technologies and Lyssn . He leads transformative university initiatives and has served in editorial roles for top journals including Computer Speech and Language and IEEE Transactions on Affective Computing .
Berrak Sisman is an Assistant Professor in the Department of Electrical and Computer Engineering at Johns Hopkins University, affiliated with the Data Science and AI Institute and the Center for Language and Speech Processing (CLSP). She leads the Speech & Machine Learning Lab (SmILe Lab), focusing on AI-driven speech technologies. She received her PhD from the National University of Singapore in 2020 and was previously a tenure-track faculty member at the University of Texas at Dallas (2022–2024). Research Interests: Her work spans artificial intelligence, speech synthesis, voice conversion, emotion analysis in speech, medical speech applications, and secure speech technology. She develops neural models for expressive and adaptive speech processing. Publications: Her recent articles (2024–2025) emphasize speech emotion recognition, zero-shot prosody control, accent conversion, and disentangled representations in TTS, reflecting a focus on cross-modal learning, robustness, and real-world applications. Awards & Grants: NSF CAREER Award (2024) Amazon Faculty Research Award (2022) Singapore Ministry of Education Award (2021) A*STAR Singapore International Graduate Award (2016–2020) Leadership: She directs the SmILe Lab, recruiting PhD/Master’s students for projects in neural speech modeling. Her grants include NSF and Amazon funding for voice conversion and emotion synthesis research.
Adam Finkelstein is a Professor in the Department of Computer Science at Princeton University, where he has been a faculty member since 1997. He holds a PhD and Master's in Computer Science from the University of Washington and a dual degree in Physics and Computer Science from Swarthmore College. Finkelstein is renowned for his interdisciplinary work at the intersection of computer graphics, audio processing, and machine learning, and he co-organized the Art of Science exhibition at Princeton. Education: PhD, Computer Science, University of Washington MS, Computer Science, University of Washington BA, Physics and Computer Science, Swarthmore College His research spans audio processing (e.g., speech enhancement, voice conversion, audio metrics), computer graphics (e.g., line drawing algorithms, stylized rendering, image manipulation), and machine learning (e.g., self-supervised learning, differentiable programming). His work often bridges technical and creative domains, exemplified by collaborations at Pixar and Adobe Creative Technologies Lab. Recent publications highlight advancements in audio super-resolution and voice conversion using deep learning frameworks, as well as stylized line rendering for animated 3D models. His contributions to perceptual audio metrics and shader optimization further underscore his impact on human-centric computational systems. Scientific Awards: NSF CAREER Award Alfred P. Sloan Fellowship Fellow of the Association for Computing Machinery (ACM) Finkelstein has secured foundational grants for his research and actively mentors students, though no specific advisees are listed. He also explores collaborative tools for internet music performance, reflecting his broader interest in distributed systems and user interfaces.
Bryan Pardo is a Professor of Computer Science at Northwestern University and head of the Interactive Audio Lab. He co-directs the Northwestern Center for Human Computer Interaction + Design and chairs the Computer Science Diversity Committee. He teaches courses in Deep Learning, Machine Learning, Generative Modeling, and Digital Music Instrument Design. PhD in Computer Science and Engineering, University of Michigan MMus in Jazz and Improvisation, University of Michigan MS in Computer Science, Ohio State University BMus in Jazz Composition, Ohio State University His research focuses on machine understanding and manipulation of sound, particularly in music and speech domains. Key areas include Machine Learning (e.g., automated gradient clipping), Signal Processing (e.g., Multi-scale Common-fate Transform), and Human Computer Interaction. Applications involve inclusive audio interfaces, audio search engines, source separation, natural language-controlled audio effects, privacy-preserving adversarial attacks on voice recognition, and music co-creation tools. Recent publications highlight advancements in neural watermarking (MaskMark), masked acoustic modeling (VampNet), and real-time adversarial privacy systems for speech. His lab's work has been applied in Adobe's AI-powered audio editor and Lexie B2 hearing aids. Scientific Awards: $1.8 million NSF Future of Work award $440K NSF grant for accessible music programming $200K Toyota grant $100K Sony grant TorchCrepe pitch tracker: 20 million+ downloads Bryan Pardo advises PhD student Max Morrison and collaborates with researchers like Patrick O'Reilly, Zeyu Jin, and Prem Seetharaman. His lab develops technologies for blind and visually impaired audio creators, including HaptEQ and Eyes-free tools.
Swati Aggarwal is a Professor in Artificial Intelligence at the Faculty of Logistics, Molde University College (HiMolde). Her research focuses on AI applications in healthcare, ethics, cognitive development, and neural networks. She holds a PhD in Neutrosophic Neural Networks, a Master's in Information Technology, and a Bachelor's in Computer Science and Engineering. Previously, she was a Marie Curie Postdoc Fellow at NTNU, working on AI models for cognitive assessment in infants (AIM_COACH project). Research Interests - AI in Health/Medicine - Ethics in AI and Societal Impact - EEG/BCI for Cognitive Assessment - Machine Learning and Deep Learning Publications Her recent work spans AI ethics, BCI applications, adversarial attacks, and healthcare diagnostics. Notable contributions include EEG-based infant perceptual monitoring (2025) and malaria detection via EfficientNet (2023). She also explores cross-lingual adversarial robustness and blockchain in hospitality systems. Labs/Teams - ABC-AI: Applied, Basic, and Conscientious AI Group - Virtual Technologies and Learning Research Group
Dongwook Yoon is an Associate Professor at the Department of Computer Science , University of British Columbia , and serves as Director of the SOCIUS Lab . He actively contributes to research in Human-Computer Interaction, Human-AI Interaction, and Virtual/Augmented Reality as a member of the Designing for People (DFP) and CAIDA research clusters. Education : PhD in Computer Science from Cornell University (2017), MS (2009) and BS (2007) in Computer Science from Seoul National University Research Focus : Designing socio-technical systems that bridge the gap between technology and human social processes, with innovations in AR/VR, multimodal interaction, and inclusive design Article Trends show his work spans: Temporal and bichronous learning environments AI self-clones and ethical implications Income inequality in virtual platforms Enhanced multimodal collaboration in VR Eyes-reduced interfaces for situational impairments Speculative participatory design for gig economy challenges Scientific Awards include: Google Academic Research Award (2024) Best Paper Award at CHI 2024 High Impact Award in Educational Technology (2024) CHCCS/SCDHM Graphics Interface Early Career Award (2023) Multiple Honorable Mentions at CHI, DIS, and CSCW Students & Collaborators range from active PhD candidates (Anika Sayara, Yuri Kim) to notable alumni (Thitaree Tanprasert, Ashish Chopra) across his SOCIUS Lab projects. His research receives funding from NSERC , KIST , Adobe , Microsoft , and Google grants.
Professor Moncef Gabbouj is a distinguished academic and researcher currently serving as Professor of Signal Processing at the Department of Computing Sciences, Faculty of Information Technology and Communication Sciences, Tampere University, Finland. Previously, he held the same position at Tampere University of Technology before the merger in 2019. He has also held visiting professorships at prestigious institutions including Hong Kong University of Technology and Science, University of Southern California, and Purdue University. Ph.D. and MSc. in Electrical Engineering from Purdue University, USA (1989 and 1986) B.Sc. in Electrical Engineering from Oklahoma State University, USA (1985) Prof. Gabbouj's research spans multiple domains within signal and image processing, with a strong focus on machine learning applications. His primary research interests include artificial intelligence, machine learning, Big Data analytics, multimedia content-based analysis, indexing and retrieval, nonlinear signal and image processing, voice conversion, and video processing and coding. His work bridges theoretical advancements with practical applications across various industries, particularly in multimedia communications and biomedical applications. His extensive publication record demonstrates a clear evolution from traditional signal processing techniques toward more sophisticated machine learning and deep learning approaches. Recent work shows increasing focus on convolutional neural networks for various applications including ECG classification, video processing, financial time-series analysis, and image recognition tasks, reflecting the broader trend in the field toward deep learning methodologies while maintaining strong foundations in signal processing theory. IEEE Fellow (2011) Member, Finnish Academy of Science and Letters (2014) Knight, First Class, of the Order of the White Rose of Finland (2006) Nokia Foundation Recognition Award (2005) Nokia Foundation Visiting Professor Award (2012) Finnish Cultural Foundation for Art and Science Award (2017) TUT Foundation Grand Award (2015) Prof. Gabbouj has supervised 64 doctoral and 72 Master's theses, demonstrating his significant contribution to academic mentoring. His research has been supported by substantial funding, including research grants totaling 8.5 million Euro (2001-2015). He has served as Academy of Finland Professor during 2011-2015 and has been involved in numerous EU research projects including Horizon, ESPRIT, HCM, IST, COST, Tempus and Erasmus programs. As Editor, Guest Editor or member of the Editorial Board of 6 international scientific journals, he has significantly influenced the academic discourse in his field. He leads the Signal Analysis and Machine Intelligence (SAMI) research group at Tampere University and serves as the Finland Site Director of the NSF IUCRC funded Center for Visual and Decision Informatics. His research unit focuses on applying advanced machine learning techniques to solve complex problems in signal processing, computer vision, and multimedia analytics, with applications ranging from healthcare to multimedia communications and financial analysis.
University of North Carolina at Chapel HillUnited States
Junier Oliva is an Assistant Professor in the Department of Computer Science at the University of North Carolina at Chapel Hill and Lead Faculty of the Master of Applied Data Science program. His research focuses on machine learning, artificial intelligence, and nonparametric statistics, particularly in high-dimensional density estimation, sequential modeling, and learning from complex/structured data. He holds a B.S., M.S., and Ph.D. in Computer Science from Carnegie Mellon University, with prior industry experience at Yahoo! and Uber ATG. Research Interests: Machine learning, artificial intelligence, nonparametric statistics, deep learning, statistical data mining, signal processing, kernel methods, and scalability. His work bridges machine and human learning via collective approaches, emphasizing simple yet flexible models for massive datasets. Awards/Grants: $592K AIM-AHEAD/NIH Grant for Human+AI Collaboration $594K NSF Grant for Scientific Discovery $500K NSF Grant for 'Machine Detectives' Project ACM BCB Best Paper Award (2022) for transparent single-cell classification work Labs/Teams: Director of the LUPA Lab, which develops machine learning techniques for holistic data understanding across domains like healthcare, earth science, and computer vision.
Gustav Eje Henter is an Assistant Professor at KTH Royal Institute of Technology, holding roles as the Head of Research at Motorica AB and a Core Team Member of the Wallenberg Research Arena (WARA) for Media and Language. He is the Secretary of the ISCA SynSIG (Special Interest Group on Speech Synthesis) and a Co-Organiser of the GENEA Workshops on Embodied Agents' Non-Verbal Behavior. His research focuses on speech synthesis, gesture generation, and multimodal interaction, with contributions to TTS systems, neural networks, and embodied AI. He leads Digital Futures, a cross-disciplinary research center addressing societal challenges through digital technologies. This center is a collaboration between KTH, Stockholm University, and RISE. His work spans foundational research to industrial applications, emphasizing ethical AI, privacy in voice conversion, and human-robot interaction. Key research themes include causal reasoning in LLMs, adversarial privacy techniques, and benchmarking frameworks like the GENEA Leaderboard. He has organized international workshops (GENEA 2021-2024) and contributed to standards in TTS evaluation methodologies. His technical innovations include HiFi-Glot for formant synthesis and Matcha-TTS for fast waveform generation. His research integrates audio, gesture, and motion synthesis with deep learning, addressing challenges in spontaneous speech synthesis, multimodal coherence, and listener perception. He advocates for rigorous evaluation practices and open challenges to advance the field's reproducibility and real-world applicability.
Lee M. Miller is a Professor and Vice Chair of Academic Affairs in the Department of Neurobiology, Physiology and Behavior at the University of California, Davis, affiliated with the Center for Mind and Brain. His research focuses on neuroengineering, computational neuroscience, and neural mechanisms underlying attention, speech processing, and multisensory integration. Research interests include the development of neural prosthetics, decoding of neuromuscular signals for prosthetic control, and understanding how auditory and visual systems interact during speech perception and attentional processes. His work bridges clinical applications (e.g., cochlear implants) with fundamental neuroscience, leveraging tools like electrophysiological recordings, EEG/MEG, and advanced signal processing techniques. Recent publications highlight innovations in electromyographic speech neuroprosthetics, the topology of neuromuscular signals, and the neural basis of speech-in-noise processing. Miller’s studies emphasize translational potential, such as improving speech synthesis from brain signals and designing haptic feedback systems for motor coordination. His contributions have advanced understanding of neural mechanisms in sensory integration, auditory attention, and the impact of cognitive factors on perception. Miller maintains a lab dedicated to these interdisciplinary efforts, with a focus on both basic science and clinical applications.
Ian Pitt is a Lecturer in Usability Engineering and Interactive Media at University College Cork (UCC). He leads the Interaction Design, E-Learning and Speech (IDEAS) Research Group, focusing on multimodal human-computer interaction, auditory interfaces, and accessibility solutions for visually impaired users. Pitt holds a D.Phil from the University of York, followed by research fellowships at Otto-von-Guericke University in Germany before joining UCC in 1997. His research interests include speech-based interfaces, e-learning systems, and accessibility technologies for blind users. Key projects include the EU-funded ENABLE Network (2011–2014) and prototype development for UniWink. He has secured significant grants, including €72,009 from IRCSET for voice analysis research and €19,478 from the EU for ICT-supported learning initiatives. Pitt has advised numerous PhD students, including Flaithri Neff (2011), Emma-Kate Crowley (2014), and current candidates Aine Kearns and Patrick Egan. His publications span journals like International Journal of Game-Based Learning and conferences such as ICCHP and ACM SIGACCESS. He has contributed to committees for conferences like CHI and the Irish HCI conference. Teaching modules include Usability Engineering, Human-Computer Interaction, and Digital Media Development. His work emphasizes inclusive design principles, with projects addressing navigation systems for blind students and adaptive e-learning frameworks. Recent research trends focus on ICT-delivered aphasia rehabilitation, emotional BCI interfaces, and multimodal learning systems. Collaborations include international partners through EU grants, reflecting his global impact in accessibility and educational technology.
Prof. Dr. Volker Dellwo is an Associate Professor of Phonetics and head of the Department of Computational Linguistics. His research focuses on phonetics, speech recognition, computational linguistics, and dialectology, with applications in forensic analysis, voice biometrics, and multimodal emotion recognition. Academic Rank: Associate Professor Department: Computational Linguistics His work explores phonetic convergence , speaker discrimination, and the role of prosodic features in voice recognition. Recent studies analyze whispered speech processing, cross-dialect accommodation, and neural mechanisms of speaker identity encoding. Key article trends include self-supervised learning for speech recognition , multimodal emotion detection , and forensic voice analysis . Subfields span acoustic variability, temporal envelope dynamics, and voice quality metrics. Publications emphasize computational phonetics , cross-linguistic studies , and neural network applications in speaker identification. Research also addresses challenges in forensic audio analysis and synthetic speech dataset generation.