Simon King is a Professor of Speech Processing at the University of Edinburgh , affiliated with the School of Philosophy, Psychology and Language Sciences . He serves as Director of the Centre for Speech Technology Research (CSTR) and teaches courses like Speech Processing and Speech Synthesis , while directing the MSc in Speech and Language Processing . Research Interests His research focuses on: Developing new acoustic models (e.g., Linear Dynamical Models, factorial-HMMs) for speech recognition Advancing unit selection and HMM-based speech synthesis Integrating articulatory measurement data for enhanced modeling Exploring perceptual measures in synthesis criteria Building multilingual speech systems to identify universal speech building blocks Publication Trends Simon's recent work emphasizes deep learning (DNNs, LSTMs) in speech synthesis, multilingual frameworks , and articulatory-acoustic feature integration . His studies often bridge grapheme-based modeling , perceptual error reduction , and noise-robust synthesis . Scientific Awards EPSRC Advanced Research Fellowship (2005-2009) Students & Collaborations He has supervised numerous PhD students including Rasmus Dall, Tom Merritt, and Srikanth Ronanki. Current research fellows like Mirjam Wester and Zhizheng Wu contribute to projects such as Natural Speech Technology (NST) and Simple4All .
Catherine Lai is a Reader (~Associate Professor) in the Department of Linguistics and English Language at the University of Edinburgh, with strong affiliations to the Centre for Speech Technology Research (CSTR) and the Institute for Language, Cognition and Computation (ILCC) in the School of Informatics. She is based in the School of Philosophy, Psychology and Language Sciences and is actively involved in research, teaching, and academic service. Department: Department of Linguistics and English Language School: School of Philosophy, Psychology and Language Sciences Research Institutes: Centre for Speech Technology Research, Institute for Language, Cognition and Computation Email: C.Lai@ed.ac.uk Her research centers on the role of prosody—non-lexical aspects of speech—in spoken communication. She investigates how prosody contributes to discourse structure, information structure, and affect in dialogue, using interdisciplinary methods from linguistics and machine learning. Her work bridges theoretical linguistics and practical speech technology, aiming to improve spoken language understanding and synthesis systems. She is particularly interested in how prosody shapes listener expectations and how affect and topic are expressed and perceived in conversation. Her recent publications reflect a strong focus on self-supervised learning in speech models, emotion recognition, ASR error correction using large language models, cognitive state classification, and ethical considerations in language technology. She explores topics such as the uncanny valley in synthetic speech, gender expression through voice, and community-centered development of language technologies. Prize from Scopus Profile Catherine Lai has supervised several PhD students, including Leimin Tian and Yuanchao Li, and has been involved in significant research projects, such as a Toyota-funded initiative on spoken dialogue for robot companions. She has secured multiple grants and leads a research agenda that integrates theoretical inquiry with real-world applications in assistive technologies and social science. Her academic service includes organizing major conferences like Interspeech and UK and Ireland Speech. She is a key member of research teams at CSTR and ILCC, collaborating across disciplines to advance the understanding of spoken communication and the development of robust, ethical speech technologies.
Mark Gales is Professor of Information Engineering at the University of Cambridge and an Official Fellow at Emmanuel College. He is currently on sabbatical leave for the 2024/25 academic year. Prior to his academic career, he worked as a consultant at Roke Manor Research Ltd, developing radar systems, before transitioning to speech and language processing. PhD in 'Model-Based Techniques for Robust Speech Recognition' (University of Cambridge, 1995) BA in Electrical and Information Sciences (University of Cambridge, 1988) His research focuses on speech and language processing , particularly in automated language assessment and low-resource speech technology . He leads the Automated Language Teaching and Assessment (ALTA) Institute , which collaborates with Cambridge University Press & Assessment (CUP&A) to develop commercial tools like Linguaskill and Speak & Improve . These platforms provide automated spoken/written assessment for millions of users globally. Recent publications highlight his work in LLM-driven speech processing , including adversarial attacks on foundation models, end-to-end spoken error correction, and uncertainty estimation frameworks. His team's research spans multilingual capabilities, with deployments in languages ranging from Dholuo to Tok Pisin . Awards : IEEE Fellow, ISCA Fellow Leadership : Fellows' Steward at Emmanuel College Mark has contributed extensively to Hidden Markov Model (HMM) applications in speech recognition, which underpinned early automatic speech systems. His work now bridges LLM-based language assessment with cross-lingual transfer learning and robustness testing for real-world deployments.
Professor Maja Pantic is a Professor of Affective & Behavioural Computing at the Department of Computing, Faculty of Engineering, Imperial College London. Her research focuses on artificial intelligence, image processing, and audio-visual speech recognition. She leads projects in multimodal systems, including facial analysis, emotion recognition, and speech-driven animation. Affiliations include the AI for Healthcare initiative, the Artificial Intelligence Network, and the Machine Learning Network. Her work addresses challenges in real-time speech enhancement, cross-modal learning, and synthetic data generation. Recent publications emphasize advancements in audiovisual speech synthesis, lip-reading, and emotion-aware systems. She has contributed to datasets like KAN-AV and SEWA DB, advancing research in face analysis and affective computing.
Dr. Armin Mustafa is an Associate Professor in Computer Vision and AI at the University of Surrey, where he holds a prestigious Royal Academy of Engineering Research Fellow position. He is affiliated with the Centre for Vision, Speech and Signal Processing (CVSSP), the School of Computer Science and Electronic Engineering, and the Surrey Institute for People-Centred Artificial Intelligence (PAI). His research focuses on developing AI systems for visual understanding of complex dynamic scenes, with applications in entertainment, autonomous systems, and augmented/virtual reality. Dr. Mustafa completed his PhD in general dynamic scene reconstruction from multi-view videos in 2016 from the University of Surrey under the supervision of Prof. Adrian Hilton. Prior to his doctoral studies, he worked for three years (2010-2013) at Samsung Research Institute in Bangalore, India, in the field of Computer Vision. His research expertise spans Computer Vision, Scene Understanding, 3D/4D Vision, Virtual Reality, Light Fields, Machine Learning, Video Captioning, Augmented Reality, Artificial Intelligence, and Audio-visual Video Understanding. Dr. Mustafa has pioneered advances in 4D vision, NLP, and Scene Understanding over the past decade, with a particular focus on enabling machines to model and interpret real-world environments for socially beneficial applications. His work bridges theoretical advances in computer vision with practical applications in media production, virtual reality, and autonomous systems. Analysis of Dr. Mustafa's recent publications reveals a strong focus on multimodal learning, particularly the integration of audio and visual information for scene understanding. His work spans diverse areas including shadow detection and removal, audio event classification, video captioning, person image generation, and dynamic scene reconstruction. A notable trend is his exploration of transformer architectures for both vision and audio tasks, as well as the application of self-supervised learning techniques to reduce dependency on labeled data. Dr. Mustafa has received numerous prestigious awards: 2018 - Research Fellowship, The Royal Academy of Engineering, UK 2017 - Young Researcher award, CVPR 2016 - Doctoral Consortium grant, CVPR 2015 - BMVA travel grant for ICCV 2014 - Set-Squared Research to Innovator grant 2013 - Overseas Research Scholarship, FEPS, The University of Surrey 2010 - Cadence Silver Medal, Indian Institute of Technology, Kanpur As a dedicated mentor, Dr. Mustafa supervises several PhD students working on cutting-edge topics including multi-person reconstruction, audio-visual scene understanding, and automatic storyboard generation. His research is supported by significant grants including a £15 million UKRI Prosperity Partnership with the BBC (AI4ME), a 5-year Royal Academy of Engineering fellowship (4D Vision for Perceptive Machines), and multiple projects with industry partners such as Figment Productions and Foundry. Dr. Mustafa is an active member of the Centre for Vision, Speech and Signal Processing (CVSSP), one of the world's leading research centers in vision, speech, and signal processing. He also contributes to the Surrey Institute for People-Centred Artificial Intelligence (PAI), where he serves as a Surrey AI Fellow. His work often involves collaboration with industry partners and other academic institutions across Europe.
Professor Stefan Bleeck is a Professor of Hearing Science and Technology at the University of Southampton, leading the Hearing and Balance Centre and directing the Institute of Sound and Vibration Research (ISVR). His research focuses on the intersection of hearing science, audiology, and signal processing, with specialties in bio-inspired auditory modeling, speech intelligibility in noise, cochlear implants, and auditory evoked potentials. He holds a PhD in computational neuroscience and has held roles including Head of the Hearing and Balance Centre. Awards include Vice-Chancellor's Teaching Awards (2009) and Google Research Awards (2012). Education: Diploma in Physics (University of Darmstadt, 1995), PhD in Computational Neuroscience (University of Darmstadt, 2000). Research spans experimental, computational, and clinical approaches to improve hearing aids and cochlear implants. Active projects include developing speech enhancement algorithms, antiphasic speech tests for hidden hearing loss, and neural-space speech processing. Supervises multiple PhD students in engineering and computer science. Publications highlight advancements in speech enhancement, bio-inspired models, and cross-linguistic hearing tests. Collaborates with institutions like Google and the European Union on projects funded by EPSRC, Cancer Research UK, and others. His work aims to enhance speech understanding for hearing-impaired individuals through innovative signal processing and auditory modeling.
Li Nguyen is an Assistant Professor of Linguistics and Multilingual Studies at Nanyang Technological University (NTU), Singapore . Their work bridges linguistics with computational approaches, focusing on language variation, contact phenomena, and multilingual NLP. Education: PhD in Linguistics (University of Cambridge), Master’s in General and Applied Linguistics (Australian National University) Research interests include: Language variation and change in multilingual/diasporic communities Computational sociolinguistics and NLP for low-resource varieties Syntax-pragmatic interface in bicultural contexts Code-switching and heritage language documentation Recent publications highlight collaborations on Vietnamese-English and other code-switched language pairs, with a focus on NLP applications and corpus-based analysis. Their work has gained recognition through grants like the NTU Start-up Grant and Cambridge Language Sciences funding . Scientific awards: Cambridge International Scholarship, Philological Society Fieldwork Grant Li Nguyen actively collaborates on projects involving sociolinguistically informed NLP and community-driven language documentation. They have contributed to the development of the CanVEC corpus for Vietnamese-English speech research.
Roles: Prof Peter Bell holds a personal chair in speech technology at the University of Edinburgh's School of Informatics and is a core member of the Centre for Speech Technology Research (CSTR). His primary research focus is automatic speech recognition (ASR), particularly in cross-domain adaptation, lightly supervised training, and minority language systems. He teaches the Automatic Speech Recognition course and advises multiple PhD students. Research Interests: Prof Bell's work spans ASR system development for diverse domains, audio-visual integration, end-to-end models, and under-resourced languages. His projects include the CoG-MHEAR healthcare initiative and the Unmute project addressing language marginalization. He has pioneered techniques for speaker adaptation, raw-waveform modeling, and multi-task learning. Commercial Activities: He advises industry on speech tech adoption, co-founded Quorate Technology (acquired by LSEG), and provides consultancy to firms developing speech solutions. His work bridges academic research with commercial impact through projects like the BBC's MGB Challenge and EU-funded SUMMA platform. Grants & Projects: Leads EPSRC-funded CoG-MHEAR and Unmute initiatives, collaborates on IARPA MATERIAL for low-resource ASR, and contributed to the SpeechWave waveform-based ASR project. His research has been supported by Bloomberg, Ericsson, Samsung, and Toshiba. Labs & Teams: Active in CSTR, leading teams in speech representation learning, adaptation techniques, and multi-modal ASR. His lab supports interdisciplinary work with NLP, HCI, and biomedical engineering groups. Personal: A passionate hillwalker, he explores Scottish Highlands and Corbetts. Previously active in Edinburgh University Hillwalking Club, his outdoor pursuits reflect his disciplined approach to research exploration.
Dr Marieke Meelen is an Associate Professor in Historical Linguistics at the University of Cambridge, affiliated with the Faculty of Modern and Medieval Languages and Linguistics. She serves as a Fellow, Tutor, and Director of Studies at Trinity Hall. Research Interests: Historical Linguistics, Comparative Syntax, Information Structure, Grammaticalisation & Pragmaticalisation, NLP for low-resource languages, Celtic & Tibeto-Burman languages Projects: PI for ELDP-funded endangered language documentation in Nepal; collaborator on ERC-funded 'PaganTibet' and AHRC-funded 'Emergence of Egophoricity' Key Contributions: Development of ASR and HTR models for Tibeto-Burman languages; historical treebank of Welsh; computational approaches to Celtic and Tibeto-Burman syntax Publications & Grants focus on syntactic reconstruction, corpus creation, and NLP for endangered languages. She mentors PhD students working on Celtic or Tibeto-Burman languages and consults for computational linguistics projects.
Lu Yin is an Assistant Professor in the School of Computer Science and Electronic Engineering at the University of Surrey. He holds affiliations as a long-term visiting researcher at Eindhoven University of Technology (TU/e) and collaborator with the Visual Informatics Group (VITA) at the University of Texas at Austin. Previously, he served as a Postdoctoral Fellow at TU/e and worked as a research scientist intern at Google's New York City office. His work bridges academic and industrial research, focusing on AI Efficiency, AI for Science, and Large Language Models. His research emphasizes optimizing neural networks through sparsity techniques, including pruning strategies for LLMs and vision models. Notable contributions include the OWL method for LLM pruning and Lottery Pools for improving sparse network performance. Yin actively collaborates with institutions like TU/e, Google Research, and Intel Research, and has organized conferences such as CAPBS 2025 and CAI 2025 Workshops. Yin has secured significant grants, including a 10,000,000 NWO-funded grant for NVIDIA A100 GPU resources. He has delivered invited talks at prestigious institutions like Carnegie Mellon University and City University of Hong Kong. His work has been recognized with the Best Paper Award from LoG 2022.
Josh Reiss is a Professor of Audio Engineering at Queen Mary University of London (QMUL), part of the School of Electronic Engineering and Computer Science . He holds additional roles including President-Elect and Fellow of the Audio Engineering Society (AES), and Visiting Professor at Birmingham City University. His research focuses on audio signal processing, procedural audio, and intelligent music production. He earned degrees including BSc in Physics, BSc in Mathematics, and a PhD. Research & Awards : Reiss has published over 200 papers, authored books like Intelligent Music Production , and received awards such as the AES Board of Governors Award (2009, 2010) and Best JAES Paper 2016. His work spans sound synthesis, dynamic range compression, and live audio systems. Teaching & Industry : Teaches modules like Artificial Intelligence and Sound Design. Co-founded startups LandR (AI mixing), Tonz, and Nemisindo. Leads the Centre for Digital Music at QMUL, advancing research in audio technology. Labs & Teams : Active in the Centre for Digital Music, collaborating on projects like the Open Multitrack Testbed and semantic audio evaluation tools.
Ajitha Rajan is a Professor (Personal Chair of Software Testing and Verification) at the School of Informatics, University of Edinburgh. Previously, she was a post-doctoral researcher at Oxford University's Computer Science Department and Laboratoire d'Informatique de Grenoble (LIG) in France. She earned her PhD in Computer Science from the University of Minnesota in August 2009 under the supervision of Prof. Mats Heimdahl. Her research spans two main directions: Automated Software Testing Techniques covering test input generation, test oracles, and coverage metrics; and Biomedical Artificial Intelligence focusing on cancer survival models, interpretability for biological sequences, and medical images. Her work bridges software engineering and biomedical applications, particularly in the development of trustworthy AI systems for healthcare. Professor Rajan leads several significant research projects including a Royal Society Industry Fellowship (2022-2025) on AutoTest for autonomous vehicle perception safety, the H2020 European Project KATY (2021-2025) on AI for genomics and personalized medicine where she serves as Edinburgh Lead PI, and an EPSRC Trustworthy Autonomous Systems Node project (2020-2024). Her recent publications demonstrate strong activity across software testing, explainable AI, and biomedical applications, with numerous papers accepted to top conferences in 2025 including ML for Healthcare, IJCAI, and ESEM. Among her scientific recognitions, she received the Best Reviewer Award at ISSTA'25. Her work has been consistently published in leading venues including ICSE, ICASSP, and Communications Biology. Professor Rajan actively mentors PhD students working on diverse topics from automated testing of speech recognition systems to explainable AI for medical image analysis and cancer immunotherapy. She teaches undergraduate courses in Software Testing, Computer Programming, and Embedded Systems, and has been instrumental in establishing several fully funded PhD positions through Centres for Doctoral Training at the University of Edinburgh.
Dr. Mo El-Haj is a Reader (Associate Professor) in Natural Language Processing (NLP) at the College of Engineering & Computer Science, VinUniversity, Hanoi, Vietnam, and holds a visiting role at Lancaster University. He specializes in NLP with a focus on Financial NLP, Arabic NLP, and multilingual systems for under-resourced languages. Education: PhD in Computer Science (University of Essex, 2012), MSc in Information Systems (University of Jordan, 2008), BSc in Computer Information Systems (University of Jordan, 2005). Awards include the FHEA Fellowship (2021) and the 2016 BBC NewsHack Best Tool award. Research interests include text summarization, financial narrative processing, biomedical NLP, and corpus linguistics. He leads the VinNLP research group and has supervised/co-supervised over 20 PhD students. Notable projects include the Welsh Automatic Text Summarisation tool (ACC) and the FreeTxt bilingual analysis toolkit, funded by the Welsh Government and AHRC. Publications span 92+ works in top journals/conferences like Computational Linguistics and LREC. Active in organizing workshops (e.g., WACL-4, FinNLP) and has served as external/internal PhD examiner at UK universities.
Dr. Ruchit Agrawal is an Assistant Professor of Computer Science and Head of Computer Science Outreach at the University of Birmingham Dubai. Previously, he served as a Postdoctoral Researcher in AI for Healthcare at the University of Oxford’s Computational Health Informatics Lab, and as a Marie Curie AI Researcher in the transnational MIP-Frontiers project at Queen Mary University of London. His work focuses on optimizing healthcare systems using Machine Learning, alongside contributions to Natural Language Processing, Audio Signal Processing, and Multimodal Deep Learning. He holds a PhD in Computer Science from Queen Mary University of London and an MS by Research from IIIT Hyderabad. Education: PhD in Computer Science (Queen Mary University of London, 2021) MS by Research in Machine Translation (IIIT Hyderabad, 2017) Research Interests: Clinical Machine Learning for healthcare system optimization Natural Language Processing with a focus on Indian languages and context-aware models Audio Signal Processing for music performance analysis and stuttering detection Development of multimodal deep learning frameworks for diverse applications Adaptive AI systems leveraging contextual and positional encoding techniques Publications highlight trends in healthcare AI, multilingual NLP, and audio-visual alignment. Recent work includes Arabic sentiment analysis, stuttering detection via MMSD-Net, and stock price prediction using FB-GAN. Earlier contributions address structure-aware synchronization in music performance data and transformer-based post-editing for low-resource languages. His research bridges theoretical advancements with practical implementations in clinical, financial, and cross-modal domains. Scientific awards include the prestigious Marie Skłodowska-Curie scholarship (2017–2020) supporting his deep learning research in audio signal processing. Advising and grants: While no formal advisees are listed, his roles involve leading outreach initiatives and guiding collaborative projects at the Computational Health Informatics Lab during his postdoctoral tenure. Labs/Teams: Active member of the Computational Health Informatics Lab (Oxford) and Machine Translation group at FBK (Italy). His work also intersects with the MIP-Frontiers transnational research project.
Dr. Anton Ragni is a Senior Lecturer in Speech and Language Technologies at the University of Sheffield's School of Computer Science, where he serves as Assessments Lead and contributes to the Speech and Hearing (SpandH) research group. His educational background includes: BEng in Information Technology from the University of Tartu (2005) MEng in Information Technology from the University of Tartu (2007) PhD from the University of Cambridge (2013) Ragni's research centers on machine learning approaches for speech and language processing, with core expertise in automatic speech recognition (ASR), expressive speech synthesis, spoken language translation, information retrieval, and conversation modeling. His work increasingly integrates self-supervised learning and foundation models to address challenges in speech technology and cross-domain applications like music processing. Analysis of his recent publications reveals a strong trend toward applying speech processing techniques to music understanding and developing robust ASR systems for specialized populations, including hearing-impaired users and children. His work demonstrates consistent innovation in leveraging contextual information and novel architectures like energy-based models. His scientific recognition includes: Best Student Paper Award at IEEE ASRU 2011 for 'Generative kernels for noise robust ASR' Ragni has secured significant research funding as Principal Investigator and Co-Principal Investigator: EPSRC grant 'Exemplar-based Expressive Speech Synthesis' (2021-2023, £218,290) as PI Innovate UK grant 'Automatic voice conversion for transforming professional adult voice actors to artificial child voice actors' (2021-2023, £173,605) as Co-PI He actively contributes to the Speech and Hearing research group, focusing on advancing speech technology through interdisciplinary collaboration and real-world applications.