Chris Donahue is an Assistant Professor in the Computer Science Department at Carnegie Mellon University . He also serves as a part-time Research Scientist at Google DeepMind on the Magenta team. His work focuses on leveraging generative AI to enhance human creativity, particularly in music. Education: PhD in Computer Science (UC San Diego), Postdoctoral Scholar (Stanford University) His research spans controllable generative modeling of music and audio , with a focus on real-time interactive systems. Projects like Piano Genie , Beat Sage , and Copilot Arena demonstrate his commitment to real-world deployment. His Generative Creativity Lab (G-CLef) explores AI applications beyond music, including programming and natural language. Recent publications highlight advancements in multimodal music evaluation , real-time adaptation , and AI-driven sound morphing . He co-developed Magenta RealTime , an open-weight real-time music generation model, and MusicFX DJ Mode . Scientific Awards: Best Paper Award (top 1) at NAACL Student Research Workshop 2025 Best Paper Award (top 1% of submissions) at CHI 2025 Best Paper Runner-up at ISMIR 2021 He co-advises PhD students like Wayne Chi (NDSEG Fellow) and mentors Irmak Bukey . His lab receives support from the AIxArts incubator fund at CMU .
Alexander Schwing is an Associate Professor in the Department of Electrical and Computer Engineering and Computer Science at the University of Illinois at Urbana-Champaign, affiliated with the Coordinated Science Laboratory. His research focuses on machine learning and computer vision with applications in 3D scene understanding, generative modeling, and multi-agent systems. Education: Diploma in Electrical Engineering and Information Technology, Technical University of Munich (TUM) PhD in Computer Science, ETH Zurich Postdoctoral Fellow, University of Toronto Research Interests: Structured prediction in deep learning Generative adversarial networks and stability Multi-modal vision-language models 3D scene reconstruction from single images Embodied agent collaboration Semantic segmentation with temporal coherence Recent Publications: Highlight trends in neural rendering, video object segmentation, and reinforcement learning with applications to 3D modeling and multi-agent systems. Notable innovations include SAIL-VOS dataset for amodal segmentation and NeRFDeformer for single-view scene transformation. Scientific Awards: NSF CAREER Award, 3M and Amazon research awards, multiple student recognition awards, ETH Zurich PhD medal, and best paper at Intelligent Tutoring Systems 2014. Teaching: Offers graduate courses in Pattern Recognition (ECE 544) and Machine Learning (CS 446/ECE 449). Previously taught at University of Toronto and ETH Zurich. Labs & Collaborations: Leads research at Coordinated Science Laboratory (UIUC) with collaborations across University of Toronto, ETH Zurich, and industry partners like Samsung SAIT and Amazon.
Andrew Zisserman is a Royal Society Research Professor at the University of Oxford's Department of Engineering Science, affiliated with the Visual Geometry Group (VGG). His research focuses on computer vision, artificial intelligence, and neural networks, with significant contributions to multimodal learning, video understanding, and 3D scene analysis. He leads projects exploring visual-language models, audio-visual synchronization, and clinical imaging applications. Key research areas include: Video analysis and temporal modeling Multimodal systems for sign language translation and action recognition 3D shape estimation and physical property inference Foundation models and cross-modal retrieval Recent work highlights: Developed Flamingo and Tapir models for video-language tasks Advancements in spinal MRI analysis and clinical imaging Leadership in EGO4D and VoxCeleb challenges Honors include Fellowship of the Royal Society (FRS) and the ISSLS Prize in Clinical Science 2023 for spinal analysis innovations. His lab collaborates globally, emphasizing real-world applications in healthcare and autonomous systems.
Nasir Memon is a Professor of Computer Science and Engineering at the New York University Tandon School of Engineering and concurrently serves as the Dean of Engineering at NYU Shanghai. He has been a faculty member at NYU Tandon since September 1998. Memon is a co-founder of NYU's Center for Cyber Security (CCS) and NYU Abu Dhabi, and the founder of key initiatives such as the OSIRIS Lab, CyberSecurity Awareness Week (CSAW), the NYU Tandon Bridge program, and the Cyber Fellows program. His work focuses on advancing cybersecurity education and addressing systemic biases in AI-driven systems. Education: Ph.D., Computer Science, University of Nebraska Master of Science, Mathematics, Birla Institute of Technology and Science (BITS), Pilani Bachelor of Engineering, Chemical Engineering, BITS, Pilani Research Interests: Media Forensics and Authentication Biometric Security and Privacy Data Compression and Privacy-Preserving Techniques Network Security and Incident Response AI Ethics and Fairness in Machine Learning Cybersecurity Education and Workforce Development Awards and Honors: IEEE Fellow (2010) SPIE Fellow (2014) Jacobs Excellence in Education Award (2002) NSF CAREER Award (1997) Advising and Grants: Advises Ph.D. students like Anubhav Jain and Govind Mittal, and mentors Master’s students such as Rishit Dholakia. Recipient of NSF grants and funding from NYU Abu Dhabi and Indiana University Bloomington for projects like computational tools for fact-checking and AI-driven bias mitigation. Labs and Teams: Directs the OSIRIS Lab, a leading research group in cybersecurity and AI. Leads initiatives at the NYU Center for Cybersecurity (CCS) and collaborates with NYU Abu Dhabi’s Center for Cyber Security.
Yonatan Bisk is an Assistant Professor at Carnegie Mellon University (CMU) in the School of Computer Science , with dual appointments in the Language Technologies Institute and Robotics Institute . His research bridges Natural Language Processing (NLP) with robotics, focusing on grounded and embodied language understanding. Assistant Professor, Language Technologies Institute, CMU (2021–Present) Courtesy Appointment, Robotics Institute, CMU Research Themes : Language as a social codification of embodied experience Interpretable multimodal model training Human-robot collaboration frameworks Embodied question-answering systems Selected Trends : His recent publications show increasing focus on cross-modal attention mechanisms (Vid2Robot), error detection in toolchains (Tools Fail), and theory-of-mind reasoning in language agents (SOTOPIA). Multimodal integration spans vision, audio, and robotic control contexts (ANAVI). Labs & Collaborations : Founder of CLAW Lab (Connecting Language to Action and the World) Collaborations with Microsoft Research, Meta Inc, and CMU's REAL (Robotics, Embodied AI, Learning) community
Massachusetts Institute of TechnologyUnited States
James Glass is a Senior Research Scientist at the Massachusetts Institute of Technology (MIT) and heads the Spoken Language Systems Group within MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL). He is also affiliated with the Harvard-MIT Division of Health Sciences and Technology. His research spans automatic speech recognition, multimodal learning, and spoken language understanding, with applications in healthcare and video analysis. Education: SM and PhD in Electrical Engineering and Computer Science from MIT His work focuses on paralinguistic speech analysis, health markers in speech, and the intersection of speech and natural language processing. Recent trends emphasize audio-visual alignment, recursive reasoning, and AI applications in cognitive disorder diagnosis. Scientific awards include IEEE Fellow, ISCA Fellow, and Associate Editor for IEEE Transactions on Pattern Analysis and Machine Intelligence. His group explores unsupervised learning, speaker verification, and social text analysis. James leads the Spoken Language Systems Group at CSAIL, collaborating with institutions like IBM and Harvard-MIT Division of Health Sciences and Technology. His research integrates vision-language models, neural audio codecs, and self-supervised frameworks.
Rong Zheng is a Professor in the Department of Computing and Software and a member of the School of Biomedical Engineering at McMaster University, Canada. She holds a Tier-1 Canada Research Chair in Mobile Computing and serves as Acting Chair of the Computing and Software department from July to December 2025. She is also an Associate Member of the Electrical & Computer Engineering department. Education: Ph.D. in Computer Science, University of Illinois, Urbana-Champaign, USA Master of Engineering (thesis) in Electrical Engineering, Tsinghua University, Beijing, China Bachelor of Engineering in Electrical Engineering, Tsinghua University, Beijing, China Dr. Zheng's research lies at the intersection of mobile computing, wireless networking, and machine learning, with a strong focus on applications for aging populations. She directs the NSERC Smart Mobility for the Aging Population CREATE program. Her work encompasses sensor development, wireless network design, and mobile data analytics to address real-world challenges in healthcare, mobility, and data center monitoring. She has developed innovative solutions like the MacQuest campus navigation app and has captured first prize in indoor localization competitions. Her recent publications demonstrate a clear trajectory toward applying wireless sensing technologies (particularly acoustic, Wi-Fi, and mmWave) to health monitoring and mobility assessment for older adults. There's a strong emphasis on developing efficient edge computing solutions that can process data in real-time on resource-constrained devices, as exemplified by her TeamNet framework for collaborative inference on the edge. Her work bridges theoretical advances with practical applications that have social impact. Scientific Awards: Tier-1 Canada Research Chair in Mobile Computing US National Science Foundation CAREER Award (2006) Joseph Ip Distinguished Engineering Fellow (2015-2018) Dr. Zheng leads the Wireless System Research Group (WiSeR) at McMaster University, which has secured significant funding including a $1.65M NSERC CREATE grant for smart mobility research for older adults. Her research has been supported by multiple funding agencies including NSERC, NSF, UH GEAR, and DURIP. She actively mentors graduate students and has developed specialized courses including CAS 772 (Mobile Data Analytics) and CAS 781 (Mobility in the Aging Population). The WiSeR group conducts impactful research on communication, networking, and data analytics issues in Cyber Physical Systems, with applications spanning healthcare, smart infrastructure, and data center monitoring. Their work on data center infrastructure monitoring networks has been featured in EurekAlert and Data Center Dynamics, and they've made significant contributions to indoor localization technology.
Dr A. I. Shihab is a Senior Lecturer at Kingston University's Faculty of Engineering, Computing and the Environment, Department of Networks and Digital Media. He teaches programming languages (C++/Java), data structures, web development, and AI/machine learning. His research focuses on affective computing and machine learning applications including: Acoustic event detection in sports environments Audio signal analysis for tennis match modeling Multi-camera visual surveillance systems Medical imaging analysis using fuzzy clustering techniques Publications demonstrate expertise in combining audio/video modalities for sports analytics (tennis rallies, court-shots) and developing Markov models for sound event sequence analysis. Contact: a.shihab@kingston.ac.uk
Professor Peter F. Driessen is a faculty member in the Department of Electrical and Computer Engineering at the University of Victoria, with a cross-appointment in the School of Music. He holds a BSc and PhD from the University of Victoria and is a Professional Engineer (PEng). His research focuses on communication systems, signal processing, control, and interdisciplinary projects in computer music and wireless technologies. Key areas include audio/video signal processing, radio propagation, sound recording, and multimedia systems. He leads the University of Victoria Propagation Laboratory, which explores radio wave propagation and Amateur radio integration with engineering education. His work spans theoretical research and applied projects like ECOSat satellite systems, software-defined radio (SDR), and innovative musical instruments such as the Radio Drum. He supervises undergraduate and graduate projects in these domains through ELEC 499 courses. Notable contributions include the APEGBC Editorial Board Award for Best Paper (2002) and patents in wireless networking and signal processing. His teaching includes courses in signal analysis and electromagnetics, and he collaborates on interdisciplinary programs like the Music/Computer Science degree. Education: BSc in Electrical Engineering, University of Victoria PhD in Electrical Engineering, University of Victoria Research Interests: Audio and video signal processing for music and media Software-defined radio and Amateur radio technologies Satellite communication and ground station development Gesture-based interfaces and musical instrument design Error mitigation in streaming audio/video Optical and microwave-photonic systems Labs & Collaborations: Propagation Laboratory (radio wave research) UVic Experimental Radio Group (Amateur radio club) UVic Satellite Design Team (ECOSat projects) UVic Centre for Aerospace Research Grants & Awards: APEGBC Editorial Board Award (2002) Multiple US patents in wireless systems and signal processing
Habib Ullah is an Associate Professor in Data Science at the Norwegian University of Life Sciences (NMBU), Norway, where he conducts research at the intersection of computer vision and machine learning. He is affiliated with the Institute of Data Science under the Faculty of Science and Technology. He has previously held academic positions at COMSATS University Islamabad, Pakistan, and the University of Ha'il, Saudi Arabia, and served as a postdoctoral researcher at The Arctic University of Norway. Educational Background: PhD in Information and Communication Technology (Computer Vision), University of Trento, Italy (2011–2015) MSc in Electronics and Computer Engineering, Hanyang University, South Korea (2007–2009) BSc in Computer Systems Engineering, NWFP University of Engineering and Technology, Pakistan (2002–2006) Habib Ullah's research is primarily focused on computer vision and machine learning, with applications in aquaculture, agriculture, and human behavior analysis. He investigates underwater fish feeding sounds using audio classification, develops zero-shot learning models for recognizing unseen classes, and applies deep learning to detect stress in salmon via skin dot patterns. He also explores AI-driven controlled environment agriculture, leveraging sensors and automation for optimal crop growth. His work emphasizes practical AI solutions for real-world challenges in environmental and biological domains. The recent publications highlight a strong trend in leveraging deep learning for zero-shot and semi-supervised learning, particularly in computer vision tasks such as sea ice classification, crowd anomaly detection, and agricultural monitoring. His research spans remote sensing, biomedical signal processing, and human activity recognition, demonstrating interdisciplinary versatility. The keywords reflect a focus on robust feature representation, knowledge transfer, and model generalization. Scientific Awards and Funding: Industrial PhD grant 'Advancing Controlled Environment Agriculture AI' from The Research Council of Norway (Project number 354125, 2 million NOK, 2024) Team member (Coordinator-Participant) in the Battery Cell Assembly Twin (BatCAT) project funded by Horizon Europe (7 mEuro, 2023–2027) Development of an AI-Based Image Analysis System for Monitoring Plant Status (Funding: 1.8 mNOK, starting 2025) Habib Ullah actively supervises PhD projects and contributes to academic service through editorial and organizational roles. He has served as an Associate Editor for IEEE Access, Guest Editor for MDPI Remote Sensing, and Editor of the Springer book Machine Learning Techniques and Sensor Applications for Human Emotion, Activity Recognition, and Support (ML-SHEARS) . He has also been a Track Chair and Program Committee Member for several international conferences, reflecting his leadership in the academic community. His research is supported by significant grants and collaborative projects, indicating strong institutional and international engagement. He is involved in multiple research teams and projects, including the BatCAT project on battery manufacturing and AI applications in controlled environment agriculture with RIFT LABS AS. His lab work integrates deep learning, sensor fusion, and data analytics for environmental and biological monitoring systems.
David J. Crandall is the Luddy Professor of Computer Science at Indiana University's Luddy School of Informatics, Computing, and Engineering. He serves as Director of the Luddy Artificial Intelligence Center and leads the IU Computer Vision Lab. With joint appointments in Informatics, Cognitive Science, Data Science, and Statistics, his work spans computer vision, machine learning, and AI. He holds a Ph.D. from Cornell University and previously worked at Eastman Kodak Research Labs. His research focuses on developing statistical and machine learning methods to analyze visual information, including object recognition, human activity analysis in video, 3D reconstruction, social media mining, and computational studies of visual attention. Key applications include egocentric vision systems, social robotics for healthcare, and cross-disciplinary collaborations with developmental psychology. Recent publications demonstrate strong emphasis on egocentric video analysis (Ego4D), human-robot interaction (CHI/HRI), and explainable AI (IJCAI). Medical imaging, nanoscale security systems, and computational social science represent emerging interdisciplinary directions. His work consistently integrates deep learning with real-world applications in health, environmental monitoring, and cultural analytics. Tracy M. Sonneborn Award (2024) Distinguished Member of the ACM (2023) Luddy Professorship (2021) NSF CAREER Grant (2013) Trustees Teaching Award (2017) He has advised over 20 Ph.D. graduates, with current students working on computer vision, robotics, and AI ethics. Major grants include $20M for the NSF AI Institute on Engaged Learning, $4.4M for trusted AI research, and funding from NIH, Google, ONR, and NASA. He directs the Computer Vision Lab and collaborates with Selma Sabanović's robotics group on social agents for older adults.
Devi Parikh is an Associate Professor at the School of Interactive Computing, Georgia Institute of Technology, and a Research Director at Meta’s FAIR lab. Her research focuses on generative models, AI for creativity, computer vision, and natural language processing. Education: B.S. in Electrical and Computer Engineering from Rowan University (2005), M.S. and Ph.D. in Electrical and Computer Engineering from Carnegie Mellon University (2007, 2009). Research interests include embodied AI, human-AI collaboration, and creative applications of AI. She has held visiting positions at Cornell, MIT, CMU, and others. Awards include NSF CAREER Award, IJCAI Computers and Thought Award, and multiple fellowships. Led development of Habitat , a platform for embodied AI research, and contributed to the Open Catalyst Project for renewable energy storage.
Shinji Watanabe is an Associate Professor at Carnegie Mellon University's Language Technologies Institute and a Courtesy Professor in the Electrical and Computer Engineering department. He holds a Ph.D. (Dr. Eng.) from Waseda University, Japan, and has held research roles at NTT Communication Science Laboratories, Mitsubishi Electric Research Laboratories (MERL), and Johns Hopkins University. His research focuses on automatic speech recognition, speech enhancement, and machine learning for speech processing. Watanabe has published over 300 peer-reviewed papers and received the Best Paper Award at IEEE ASRU 2019. His work emphasizes robust speech processing in challenging environments, multilingual models, and neural audio codecs. He leads the ESPnet toolkit development for end-to-end speech processing systems and contributes to technical committees like IEEE SLTC and APSIPA SLA. Recent research trends include streaming speech systems, universal speech enhancement (URGENT challenges), and fusion of discrete speech units with self-supervised representations. He explores scalable speech foundation models through benchmarks like ML-SUPERB 2.0 and investigates cross-modal audio-visual processing in challenges like MISP 2025. Education : B.S., M.S., Ph.D. (Waseda University) Affiliations : CMU Language Technologies Institute, CMU ECE, Former roles at MERL and Johns Hopkins Key Projects : ESPnet, OpenWhisper-Style Models, URGENT Challenge Frameworks
Adam Finkelstein is a Professor in the Department of Computer Science at Princeton University, where he has been a faculty member since 1997. He holds a PhD and Master's in Computer Science from the University of Washington and a dual degree in Physics and Computer Science from Swarthmore College. Finkelstein is renowned for his interdisciplinary work at the intersection of computer graphics, audio processing, and machine learning, and he co-organized the Art of Science exhibition at Princeton. Education: PhD, Computer Science, University of Washington MS, Computer Science, University of Washington BA, Physics and Computer Science, Swarthmore College His research spans audio processing (e.g., speech enhancement, voice conversion, audio metrics), computer graphics (e.g., line drawing algorithms, stylized rendering, image manipulation), and machine learning (e.g., self-supervised learning, differentiable programming). His work often bridges technical and creative domains, exemplified by collaborations at Pixar and Adobe Creative Technologies Lab. Recent publications highlight advancements in audio super-resolution and voice conversion using deep learning frameworks, as well as stylized line rendering for animated 3D models. His contributions to perceptual audio metrics and shader optimization further underscore his impact on human-centric computational systems. Scientific Awards: NSF CAREER Award Alfred P. Sloan Fellowship Fellow of the Association for Computing Machinery (ACM) Finkelstein has secured foundational grants for his research and actively mentors students, though no specific advisees are listed. He also explores collaborative tools for internet music performance, reflecting his broader interest in distributed systems and user interfaces.
Dr. Andrew Hines is a Researcher at the School of Computer Science, University College Dublin, specializing in machine learning applications for signal processing in speech, audio, and video domains. His work focuses on Quality of Experience (QoE) modeling, speech quality assessment, and immersive media analysis. He has held leadership roles in European COST Actions like Qualinet and CryptoAction, and previously worked in industry as a Director of Engineering. University: University College Dublin Role: Director of Research, Innovation and Impact Key Collaborations: IEEE (Senior Member), Audio Engineering Society (Ireland) Research interests center on machine learning for QoE optimization, audio-visual integration, and healthcare applications like heart sound classification and stroke rehabilitation. His recent publications explore self-supervised learning, neural speech codecs, and contextual factors in speech/audio quality assessment. Scientific contributions include awards like IEEE Senior Membership, and his work spans both academic research and industrial engineering in finance and aviation sectors. He leads the QxLab research team at UCD and develops open-source platforms such as WARP-Q and AQP for quality metrics.