Eva Zangerle is a Professor at Universität Innsbruck, Austria, with a focus on recommender systems, music information retrieval, and data science. She actively contributes to multi-method evaluation frameworks and collaborative research initiatives like PAN (Plagiarism Authorship Verification) workshops. Key research areas: Recommender Systems Evaluation, Music Emotion Recognition, Authorship Analysis Major collaborations with institutions like Zenodo, ACM, CEUR-WS.org, and RecSys conferences Her recent work explores graph neural networks for music recommendation, style change detection in multi-author texts, and cross-domain user modeling. Articles emphasize temporal modeling, multimodal data fusion, and ethical considerations in algorithmic systems. Eva leads tasks in PAN workshops and contributes to open-access datasets. She collaborates with researchers like Christine Bauer, Alan Said, and Günther Specht on improving evaluation practices in recommender systems.
Bo Wu is a Researcher at the MIT-IBM Watson AI Lab in Cambridge, MA, where he conducts pioneering research in deep learning, computer vision, natural language processing, and multimodal learning. Previously, he served as a postdoctoral research scientist at Columbia University after completing his Ph.D. at the Chinese Academy of Sciences (CAS) in Beijing, with additional research experience at Microsoft Research Asia (MSRA) and Academia Sinica. His academic foundation includes: Ph.D. in Computer Science, Chinese Academy of Sciences (CAS) Research internships at Microsoft Research Asia and Academia Sinica Wu's research focuses on advancing situated reasoning in real-world contexts, integrating neuro-symbolic approaches with deep learning for enhanced interpretability. His work spans video question answering, temporal forecasting, and multimodal understanding, with applications in social media prediction, enterprise AI, and personalized dialogue systems. He emphasizes bridging symbolic reasoning with neural networks to develop robust systems capable of handling open-world knowledge and dynamic environments. Analysis of his recent publications reveals three dominant trends: the creation of novel benchmarks for situated video reasoning (STAR, SOK-Bench), development of efficient multimodal architectures for enterprise applications (Granite Vision), and personalization techniques for language models. His research consistently merges computer vision with linguistic understanding while addressing practical constraints like real-time processing and model compression, demonstrating strong industry-academia translation. His scientific excellence is evidenced by prestigious recognitions including: IBM Master Inventor Award (2023) IBM Research Level-A Accomplishment Award (2021) ACL Best Demo Paper Award (2020) ICIP Prediction Challenge Champion (2020) Alibaba Global Vision AI Challenge Top 3 (2018) NIST TAC SM-KBP Top 1 (2019) Wu actively mentors emerging talent, currently recruiting students for vision-language projects. He provides significant academic service as Area Chair for ACM Multimedia, Senior Program Committee Member for AAAI and IJCAI, and organizer of the SMP Challenge at ACM Multimedia since 2017. His leadership extends to CVPR workshops on Multimodal Foundations Models (MMFM) and Multimodal Video Content Understanding (MVCS), while serving on program committees for NeurIPS, CVPR, ACL, and other top-tier conferences. As a core member of the MIT-IBM Watson AI Lab, Wu operates within a unique industry-academia ecosystem that fosters rapid translation of fundamental research into practical applications. His collaborative work with Chuang Gan and other researchers leverages IBM's computational resources and MIT's academic rigor, positioning him at the forefront of enterprise AI innovation where theoretical advances directly address real-world business challenges.
Florian Metze is an Adjunct Professor at the Language Technologies Institute (LTI) within the School of Computer Science at Carnegie Mellon University. His work focuses on advanced speech and audio processing, multimodal learning, and machine learning applications in under-resourced languages. He leads research in speech recognition, generative audio models, and cross-modal understanding. His research interests emphasize leveraging audiovisual data for robust speech recognition, developing scalable solutions for low-resource languages, and advancing multimodal systems through innovations like diffusion-based text-to-audio generation (Audio-Journey) and modular neural architectures (Legonn). His contributions include foundational work on wake-word detection, speaker verification (MASV), and context-aware error correction in ASR systems. Collaborative projects span speech technology for unwritten languages, child phonetic acquisition modeling, and integrating visual context into speech processing pipelines. His work often bridges theoretical advancements with practical applications in edge computing, security, and biomedical signal analysis (e.g., heart sound monitoring). Metze's research outputs include over 100 peer-reviewed articles since 2018, with recent emphasis on large language models for multitalker scenarios, efficient neural architectures, and cross-modal representation learning. He actively contributes to open-source tools like the ACLEW DiViMe diarization toolkit.
Foaad Khosmood is a Professor of Computer Science and Forbes Professor of Computer Engineering at California Polytechnic State University. As Research Director of the Institute for Advanced Technology and Public Policy (IATPP), he leads projects focusing on digital government transparency, AI applications in legislative analysis, and data science for historical research. His work spans NLP/AI, game design, and systems engineering with notable contributions to computational humor analysis, legislative data tools like Digital Democracy, and game jam methodologies documented in his Springer book Game Jams—History, Technology, and Organisation (2023). Education: PhD in Computational Linguistics from UC Santa Cruz (2011). Research interests include artificial intelligence, digital humanities, and applying NLP to policy analysis. Collaborations involve institutions like the University of Miami, Graz University of Technology, and Imperial College London. He actively develops tools for state legislative transparency, historical data recovery (e.g., African Californios project), and AI-driven journalism support systems. His recent publications emphasize explainable AI for humor analysis, legislative stance detection systems, and game-based learning frameworks. Projects like AI4Reporters and Central Coast Data Science Partnership highlight his commitment to bridging technology with civic engagement. He regularly contributes to conferences like DH, ACL, and FDG, advancing interdisciplinary applications of computational methods.
Dr. Franz J. Kurfess is a Professor in the Computer Science and Software Engineering Department at California Polytechnic State University (Cal Poly, San Luis Obispo). Since 2020, he has participated in the CSU Pre-Retirement Reduction in Time-Base (PRTB) Program, teaching primarily during Fall and Winter quarters while maintaining academic advising roles. He coordinates international exchange programs for the College of Engineering and focuses on interdisciplinary research in Artificial Intelligence (AI), Human-Computer Interaction (HCI), and Deep Learning applications in agriculture, environmental science, and public safety. Education and Professional Background: No formal education details provided in the text, but his career spans over three decades in academia with a focus on computational systems and educational innovation. His teaching includes advanced courses like Human-Computer Interaction (CSC 486) and User-Centered Design (CPE/CSC 484), emphasizing practical, project-based learning with industry collaboration. Research Interests: Combines technical expertise in AI with societal applications, including wildlife population monitoring, UAV navigation systems, agricultural automation, and ethical considerations in data science. His work bridges theory and practice through collaborations with external partners, such as the 'CSU-FORWARD' diabetes initiative and shark surveillance projects. Publications Highlight: Over 50 publications since 1983, with recent focus on AI-driven solutions for environmental challenges (e.g., livestock detection via aerial imaging) and educational technology (e.g., cross-class AI project collaboration). Early work includes foundational contributions to neural networks and knowledge management systems. Awards and Recognition: No explicit awards listed, but sustained academic contributions are evident through his prolific publication record and leadership in international exchange programs. His role as a faculty advisor for student theses and senior projects underscores his commitment to mentoring. Labs and Teams: Engaged in interdisciplinary collaborations, including the School Safety Project and UAVSim open-source development. Active in curriculum design for HCI and AI courses that integrate real-world problem-solving.
Alexandros Potamianos is an Associate Professor in the Division of Signals, Control and Robotics at the School of Electrical and Computer Engineering, National Technical University of Athens. His research focuses on Speech Processing, Natural Language Processing, and Multimodal Systems, with contributions to areas such as sentiment analysis, multimodal fusion, and robotics. Education: Diploma in Electrical and Computer Engineering, NTUA (1990) M.Sc. in Electrical Engineering, Harvard University (1991) Ph.D. in Electrical Engineering, Harvard University (1995) M.Sc. in Business Administration, New York University (2003) Research Interests: Speech and Audio Processing Machine Learning and Multimodal Fusion Robotics and Human-Computer Interaction Natural Language Processing His recent publications (2023-2025) emphasize advancements in domain adaptation for speech recognition, multimodal sentiment analysis, and regularization techniques for deep neural networks. He leads the ICCS laboratory, focusing on cognitive multimodal processing and human-robot interaction. Active in academic service, he coordinates the School's research initiatives and collaborates with international projects like BabyRobot, addressing child-robot interaction for autism spectrum disorder support.
Jian Kang is an Assistant Professor of Computer Science at the University of Rochester's Hajim School of Engineering & Applied Sciences. He joined in August 2023 after earning his Ph.D. in Computer Science from the University of Illinois at Urbana-Champaign. His research focuses on trustworthy artificial intelligence, data mining, computational social science, and algorithmic fairness, with a strong emphasis on ensuring ethical and reliable AI systems. Education: Ph.D. in Computer Science, University of Illinois at Urbana-Champaign Research Interests: Data Mining and Machine Learning Trustworthy AI and Uncertainty Quantification Algorithmic Fairness and Bias Mitigation Computational Social Science His work explores fair and reliable modeling of interconnected systems, leveraging graph learning, fairness-aware algorithms, and uncertainty analysis. Publications Trends: Recent work emphasizes fairness in graph neural networks, bias analysis in language models, and robust graph learning methods. Key themes include cross-lingual bias benchmarking, class-imbalanced learning, and theoretical foundations for fairness and uncertainty. Awards and Recognition: Rising Stars in Data Science (2022) Mavis Future Faculty Fellow (2022) Top Reviewer for LOG 2022, NeurIPS 2022, CIKM 2021, ICLR 2021, and ICML 2020 Advising and Grants: While no specific grants are listed, his research has led to impactful contributions in fairness and graph learning, with publications in top venues like ICLR, KDD, and TKDE. He advises on cutting-edge topics in trustworthy AI. Labs and Teams: His work is part of a broader effort at Rochester to advance ethical AI and computational methods, though no specific lab name is mentioned in the text.
Dr. Anjan Dutta is a Senior Lecturer in Artificial Intelligence at the University of Surrey, UK. He holds a PhD in Computer Science from the Autonomous University of Barcelona (UAB), awarded with Excellent Cum Laude and the Extraordinary PhD Thesis Award (2013-14). His research focuses on computer vision and machine learning, particularly deep multi-modal embedding, zero-shot learning, and graph neural networks. Education: PhD in Computer Science, UAB (2014) MSc in Computer Vision & AI, UAB (2010) MCA, Maulana Abul Kalam Azad University of Technology (2009) BSc Mathematics (Honours), University of Calcutta (2006) Research Interests: Deep Learning for Vision Tasks Zero-Shot and Few-Shot Learning Graph Neural Networks Multi-modal Embedding Techniques Structured Representation Learning Recent Research Trends: His work emphasizes scalable and interpretable AI systems, with notable contributions in object counting, bias reduction in neural networks, and sketch-based retrieval. Publications span top-tier venues, reflecting interdisciplinary innovation in vision and learning. Awards & Honors: Extraordinary PhD Thesis Award (UAB, 2013-14) Excellent Cum Laude PhD Award (UAB, 2014) Labs & Affiliations: Active in the Surrey Institute for People-Centred AI (PAI) and the Centre for Vision, Speech and Signal Processing (CVSSP), contributing to cross-disciplinary AI research.
Ilhan Aslan is an Associate Professor in the Department of Computer Science at Aalborg University (Denmark), affiliated with the Technical Faculty of IT and Design. His work focuses on Human-Centered Computing, with a strong emphasis on AI-driven interaction design, emotion-aware systems, and tangible/somaesthetic interfaces. He leads research in areas such as conversational AI, mobile computing, and smart environments, integrating machine learning with user-centered methodologies. Aslan's recent projects explore emotional speech processing, proactive AI agents, and creativity support tools, often bridging technical innovation with human well-being applications. His research spans over two decades, with notable contributions to mobile navigation systems (e.g., Bum Bag Navigator), gesture-based interaction, and ambient intelligence. Media coverage highlights his work on AI solutions for COPD patients and stress management applications. Aslan's approach combines rigorous technical development with ethnographic and participatory design practices, ensuring ethical and user-centric outcomes. His lab collaborates across disciplines, addressing challenges in HCI, AI ethics, and sustainable computing. Key Research Themes : Emotion-Aware AI, Human-AI Collaboration, Tangible Interaction, Mobile Health, Interactive Machine Learning Notable Projects : 'Feel my Speech' (haptic emotion conversion), 'BEHAVE AI' (ethical agent design), and somaesthetic smarthome systems Media Impact : Featured in 400,000+ Danes-relevant health tech stories and COPD patient support innovations Aslan's publications reflect a prolific trajectory from early mobile learning platforms (2000s) to cutting-edge generative AI and multimodal systems. His work consistently prioritizes real-world applicability through iterative prototyping and user studies, making him a key figure in both academic and applied HCI domains.
Vishnu Monn is an Associate Professor at Monash University Malaysia's School of IT, serving as Deputy Head of Education and Director of the Advanced Computing Platform. He holds a PhD in Engineering from Multimedia University (2016) and has over 15 years of academic and industry experience, including roles at Panasonic R&D Centre Malaysia (2005–2009) and Multimedia University (2009–2017). His research focuses on high-performance computing, predictive analytics, machine learning, computer vision, and soft robotics, with over MYR 1 million in secured grants as principal investigator. Education: B.Eng. (First-Class Honours) in Electrical and Electronics Engineering (2004) M.Eng. in Electrical and Electronics Engineering (2007) Ph.D. in Engineering (2016) Research interests include: High-performance computing architectures and applications Predictive analytics for industrial and environmental systems Machine learning models for computer vision tasks Soft robotics control systems Recent projects span autonomous driving algorithms, IoT-enabled smart cities, and soft robotics using reinforcement learning. He has published 61+ peer-reviewed articles in top journals and conferences, with recent work emphasizing multimodal learning, generative models, and blockchain optimization. Scientific contributions include leadership in interdisciplinary research, securing major grants, and establishing Monash Malaysia's high-performance computing facility. Teaching commitments include courses on big data, parallel computing, and embedded systems, with roles as chief examiner and program coordinator.
Titus Zaharia is a Professor at Telecom SudParis, affiliated with the SAMOVAR research department. His work focuses on computer vision, 3D data compression, neural networks, and augmented reality applications. He has contributed extensively to standards like MPEG-4 and MPEG-7, particularly in 3D mesh compression and multimedia indexing. Zaharia leads research in assistive technologies for visually and hearing-impaired individuals, developing systems like DEEP-HEAR and wearable devices for navigation assistance. His work bridges theoretical advancements with practical industrial applications, such as optimizing AR systems for manufacturing and improving public transportation prediction through machine learning. Zaharia has collaborated on projects like the INVENIO platform for content reuse in multimedia production, and his contributions span over 20 years of academic and applied research. Key research areas include dynamic point cloud compression, lightweight neural network libraries (e.g., FasterAI), and multimodal fusion for video advertising and accessibility. He has published over 100 peer-reviewed papers in journals like IEEE Access and Image and Vision Computing, and presented at conferences including CVPR, ICCV, and ISMAR. His research emphasizes real-world impact, addressing challenges in efficient data representation, human-centric technology, and industrial automation.
David Guy Brizan is an Associate Professor at the University of San Francisco, specializing in Natural Language Processing, Machine Learning, and Databases. His research focuses on analyzing personal and cultural/demographic information embedded in speech and typing to enhance speech recognition systems and cybersecurity measures. He holds a PhD in Computer Science from the CUNY Graduate Center, alongside an MS from San Francisco State University and a BS from Brooklyn College. Education: PhD, Computer Science, CUNY Graduate Center (2018) MS, Computer Science, San Francisco State University BS, Computer & Information Science, Brooklyn College Brizan's research explores intersections of linguistics and technology, including keystroke dynamics for user authentication, speech-based demographic prediction, and conversational style analysis. His work has applications in cybersecurity, healthcare diagnostics (e.g., Parkinson’s disease detection via speech), and political discourse analysis. He has published extensively in journals like Scientometrics and conferences such as LREC. Prior to academia, he worked as an IT Coordinator for NYC’s Department of Correction, a Software Engineer at IBM, and an Interface Developer at McKesson. His industry experience informs his research on practical machine learning solutions. His advising and grants span collaborations with institutions like the Speech Lab at Queens College, where he previously served as a research assistant. He has no listed awards but is actively involved in interdisciplinary projects linking computational methods to social and medical domains.
Alberto Abad is an Associate Professor at the Department of Computer Science and Engineering (DEI), Instituto Superior Técnico (IST), University of Lisbon, and a researcher at INESC-ID. He is the coordinator of the Human Language Technologies (HLT) laboratory at INESC-ID and the deputy coordinator of the Master in Computer Science and Engineering at IST. He holds a PhD and a degree in Telecommunication Engineering from the Technical University of Catalonia (UPC). He is also an IEEE Senior Member, ACM Member, and ISCA Member. Alberto Abad’s research interests include: Robust Speech Recognition Speaker and Language Characterization Applied Machine Learning Healthcare Applications of Speech Processing Privacy-Preserving Speech Technologies Speech as a Biomarker for Disease Detection His recent publications demonstrate a strong focus on privacy-aware speech systems, healthcare diagnostics (e.g., Parkinson’s and Alzheimer’s), children’s speech recognition, and secure speaker verification. He has contributed to over 100 peer-reviewed publications in top-tier conferences and journals such as IEEE TASLP, ICASSP, Interspeech, and LREC. Alberto Abad has been Principal Investigator for multiple national and international projects, including VITHEA (FCT), DIRHA (European), Biovisualspeech (CMU-Portugal), TAPAS ITN (EU), and currently leads the INESC-ID/IST contributions in the Accelerat.ai project. He actively collaborates with industry and startups in the field of speech technology. He has supervised over 25 Master’s theses and more than five PhD students, either as advisor or co-advisor, in areas such as speech biomarkers, privacy-preserving ML, children’s ASR, and cognitive disorder detection. He teaches graduate courses including Speech Processing, Programming Fundamentals, and Master’s Dissertation in Informatics and Computer Engineering. He is involved in the development of tools and platforms such as the SPA web-based platform for speech processing, the VITHEA virtual therapist for aphasia, and the BioVisualSpeech corpus and therapy system. His work bridges theoretical advances with real-world clinical and industrial applications.
Emmanouil Benetos is a Reader in Machine Listening and Director of Research at Queen Mary University of London's School of Electronic Engineering and Computer Science. He co-leads the School's Machine Listening Lab and is affiliated with the Centre for Digital Music, Centre for Intelligent Sensing, Digital Environment Research Institute, and Centre for Multimodal AI. His research focuses on computational audio analysis applied to music, urban sounds, and bioacoustics. Key areas include machine listening, self-supervised learning, audio representation frameworks, and multimodal AI. Current projects involve resource-efficient audio processing and large language model integration for acoustic tasks. Recent publications highlight advancements in music source separation, lyrics transcription, graph neural networks for audio, and acoustic identification systems. His work bridges machine learning with real-world applications in music technology, environmental monitoring, and audio-language models. Royal Academy of Engineering / Leverhulme Trust Research Fellow Turing Fellow at the Alan Turing Institute Royal Academy of Engineering Research Fellow As academic service, he serves as Secretary for the International Society for Music Information Retrieval (ISMIR), chair of the IEEE Technical Committee education subcommittee, associate editor for IEEE/ACM Transactions and EURASIP Journal, and Deputy Director for the UKRI Centre for Doctoral Training in Artificial Intelligence and Music (AIM).
Julian McAuley is a Professor in the Department of Computer Science at the University of California, San Diego (UCSD). His research bridges machine learning, natural language processing, and computer music, with a focus on generative models, recommender systems, and multimodal learning. He leads a lab that has produced influential datasets and frameworks for recommendation tasks. Primary Affiliation: UCSD, Department of Computer Science Research Themes: Generative AI, Recommender Systems, Music-Cognition Interfaces, Multimodal Learning McAuley's work explores the intersection of large language models (LLMs) with sequential recommendation, causal inference, and creative applications in music generation. His lab develops novel architectures like CoMMIT (multimodal instruction tuning) and SAND (LLM agent deliberation), while also advancing ethical AI through normative alignment techniques. Recent publications highlight trends in code-augmented reasoning , symbolic music processing , and contextual preference optimization . Notable applications include video-guided music synthesis, Explainable Chain-of-Thought systems, and tools for scalable self-updating models. He advises PhD students in areas spanning large language models , vision-language systems , and healthcare-driven AI . Collaborations span institutions like MIT-IBM Watson AI Lab, CMU, and companies including Google Deepmind, Meta, and Nvidia.