Prof. Leonardo Gabrielli is a Researcher at the Department of Information Engineering (Università Politecnica delle Marche). His work focuses on audio signal processing, machine learning, and embedded systems with applications in real-time audio generation, acoustic scene analysis, and smart transportation systems. He has contributed to projects such as emergency vehicle detection systems and audio enhancement techniques for reverberant environments. His research interests include neural network-based audio processing, physical modeling of musical instruments, and sustainable IoT applications. He has developed prototypes for driver-assistance systems, packet loss concealment algorithms, and biomedical signal analysis tools using deep learning approaches. Prof. Gabrielli collaborates on datasets like A3CarScene for driving scene understanding and has explored generative models for music synthesis and medical diagnostics. His work bridges theoretical signal processing with practical embedded system implementations, emphasizing real-world applications in automotive, healthcare, and environmental monitoring domains. Key Projects: Emergency vehicle detection, smart loudspeaker equalization, music streaming optimization Lab Affiliation: Department of Information Engineering's audio processing laboratories
Jacob O. Wobbrock is a Professor of Information at the University of Washington's Information School, with a courtesy appointment in the Paul G. Allen School of Computer Science & Engineering. He directs the ACE Lab, serves as Associate Director and founding Co-Director Emeritus of the CREATE center, and is a founding member of the DUB Group and MHCI+D degree program. His research focuses on human-computer interaction with emphasis on accessible and mobile computing, seeking to scientifically understand people's experiences with interactive technologies and improve those experiences through innovative design. Professor Wobbrock's educational background includes: Ph.D. in Human-Computer Interaction from Carnegie Mellon University (2006) M.S. in Computer Science from Stanford University (2000) B.S. in Symbolic Systems from Stanford University (1998) Wobbrock's research centers on scientifically understanding people's performance and experiences with interactive technologies, with particular focus on designing better interaction techniques and systems for people with disabilities. His work spans text entry, pointing, touch, and gesture interfaces; human performance measurement and modeling; mobile HCI; and virtual reality. He pioneered the concept of Ability-Based Design, which advocates for designing technologies that work well across the ability spectrum rather than creating specialized solutions only for people with disabilities. His recent publications (2023-2025) demonstrate continued innovation in accessible computing, particularly for blind and low-vision users through projects like SceneVR, A11yShape, and VizXpress. Professor Wobbrock's scientific contributions have been recognized with numerous prestigious awards: ACM Fellow (2021) CHI Academy (2019) SIGCHI Social Impact Award (2017) NSF CAREER Award (2010) Four 10-year lasting impact awards (ASSETS 2019, ICMI 2022, UIST 2024, ASSETS 2025) Wobbrock has successfully mentored numerous Ph.D. students who now hold faculty positions at top institutions including Harvard, Cornell, Colorado, Washington, Brown, and Simon Fraser. His research has been supported by significant funding, including an NSF CAREER award and seven other National Science Foundation grants. He also has industry experience as co-founder and former CEO of AnswerDash, which was acquired by CloudEngage in 2020, bringing practical insights to his academic work. Professor Wobbrock directs the ACE Lab (Advanced Computing Experiences Laboratory), which is part of the larger DUB Group, a cross-campus consortium of HCI researchers at the University of Washington. He is also Associate Director and founding Co-Director Emeritus of the CREATE center (Center for Research and Education on Accessible Technology and Experiences), which focuses on developing accessible technologies and educational opportunities in this field. These labs provide rich interdisciplinary environments for students and researchers working on cutting-edge problems in accessibility and human-computer interaction.
Zheng Ge is an active researcher specializing in computer vision and deep learning, with a prolific publication record spanning from 2018 to 2025. They have authored or co-authored 81 publications in top-tier venues including CVPR, ICLR, ECCV, and NeurIPS, demonstrating consistent research productivity and impact in the field of artificial intelligence. Dr. Ge's primary research interests focus on computer vision with particular emphasis on object detection, 3D perception, and multimodal learning systems. Their work bridges theoretical advancements with practical applications, especially in video understanding, image generation, and large vision-language models. Recent research directions show a strong focus on developing more efficient and capable multimodal systems that can handle complex visual tasks with human-like understanding. The analysis of their recent publications reveals a clear trend toward building more sophisticated multimodal systems that integrate vision, language, and temporal understanding. Their work on video generation, perception modeling, and instruction tuning demonstrates a strategic focus on advancing the capabilities of multimodal large language models while addressing practical challenges in scalability and robustness. While specific awards are not documented in the available publication records, Zheng Ge's consistent presence in top-tier computer vision and machine learning conferences indicates significant recognition within the research community. Their work has contributed to advancing state-of-the-art techniques in multiple subfields of computer vision. Dr. Ge appears to be actively collaborating with a large research team, frequently working with researchers such as Xiangyu Zhang, Jianjian Sun, Chunrui Han, and Liang Zhao. These collaborations span multiple institutions and research groups, suggesting a well-integrated position within the computer vision research ecosystem. While specific grant information isn't available in the publication metadata, the scale and scope of their projects indicate substantial research funding support. Based on publication patterns and technical focus, Zheng Ge appears to be part of a large research organization or lab specializing in computer vision and multimodal AI systems. Their work on technical reports for video foundation models and large-scale multimodal systems suggests involvement in industry-leading AI research initiatives with strong engineering capabilities.
Wenwen Zhang is a Professor at the Department of Electronic Engineering, College of Information Science and Electronic Engineering, Zhejiang University. With an extensive publication record spanning from 2016 to 2025, Dr. Zhang has established herself as a prominent researcher in multiple interdisciplinary fields at the intersection of computer vision, machine learning, and sensor systems. Her work demonstrates significant contributions to medical imaging, sensor array systems, wireless communications, and AI-assisted applications. Dr. Zhang's research interests encompass a wide range of topics including medical image analysis, sensor array systems, wireless communications, and AI-assisted applications. Her work demonstrates particular expertise in developing innovative deep learning architectures for medical imaging tasks such as cardiac segmentation and nuclei detection, as well as creating sophisticated models for gas sensing and wireless communication systems. She has made significant contributions to the fields of one-shot object detection, medical image segmentation, and sensor fusion techniques, with her research often bridging theoretical advancements with practical applications in healthcare and engineering. Analysis of Dr. Zhang's recent publications reveals a strong focus on cutting-edge deep learning approaches applied to medical imaging and sensor systems. Her work shows increasing sophistication in model architectures, moving from traditional CNNs to more complex transformer-based and hybrid models. There's a clear trajectory toward more explainable and clinically relevant AI systems, particularly in medical applications. Her research also demonstrates growing interest in multimodal approaches, combining different types of data and sensors to improve system performance. Dr. Zhang maintains active collaborations with researchers at Zhejiang University, particularly with Yuanjin Zheng and Zhiping Lin in the field of electronic engineering and sensor systems. She also collaborates extensively with Fei-Yue Wang from the University of Chinese Academy of Sciences, evidenced by multiple publications on parallel vision frameworks. Her international collaborations include work with researchers from institutions in Canada on intelligent knee sleeves and other biomedical applications. Her publication record shows consistent productivity with 12 publications in 2025 (as of this writing), 25 in 2024, and 26 in 2023, indicating an active and growing research program across multiple high-impact journals and conferences.
Subramanian Ramanathan is a researcher at the School of Computing, National College of Ireland , specializing in affective computing, multimodal behavior analysis, and human-computer interaction. His work spans machine learning, computer vision, and neuro-signal processing. Key research areas: Affective Computing, Deep Learning, Stress Detection, Depression Biomarkers Notable collaborations: Roland Göcke, Abhinav Dhall, Nicu Sebe His publications focus on: EEG-based cognitive load estimation Head motion pattern analysis for mental health Deepfake detection systems Audio-visual saliency prediction Transformers in behavioral modeling Stress detection via multimodal fusion Recent work includes medical imaging applications for autism detection and computational advertising systems using emotion recognition. He contributes to open science through dataset creation (SALSA, DECAF) and collaborative research in affective computing.
Alexander Schindler is an External Lecturer at the Department of Information Systems Engineering at Technische Universität Wien. His research focuses on audio signal processing, music information retrieval, and deep learning applications in multimodal analysis. He coordinates the approacH project (2010–2013) funded by the European Commission, exploring audio-visual search engines. His academic background includes a Dipl.-Ing. in Technical Engineering and a Dr.techn. from TU Wien. Key research areas include acoustic scene classification, deepfake detection, and music video analysis. He has supervised students on topics like bird song identification and machine outage prediction. Notable publications span from unsupervised cross-modal learning (2020) to multi-modal MIR frameworks (2019). He holds a Bakk.techn. and advanced technical qualifications. Recent work includes advancements in deepfake audio detection (2025) and audio-visual surveillance systems (2024). His projects bridge theoretical research with practical applications in forensic analysis and industrial predictive maintenance. He contributes to international conferences like ACM SAC and DCASE, focusing on neural network architectures for audio analysis.
Anina Rich is a Professor at Macquarie University's School of Psychological Sciences and an ARC Future Fellow. She is affiliated with the Performance and Expertise Research Centre and the Perception in Action Research Centre. Her research focuses on sensory processing, particularly selective attention mechanisms and multisensory integration, including synaesthesia studies. She has received numerous awards, including being a finalist in the Australian Museum Eureka Prizes (2001, 2002) and the Australian Postgraduate Award (2001-2004). Her work combines psychophysics, neuroimaging, and cognitive neuroscience to explore attention dynamics and synaesthesia. Research Interests: - Selective attention: Balancing voluntary and involuntary attention, applied in medical imaging contexts, sustained attention, and neural underpinnings. - Multisensory integration: Investigating synaesthesia (e.g., mirror-touch, grapheme-color), exploring how sensory modalities interact. - Neuroimaging techniques: Using fMRI and MEG to study brain regions involved in attention and perception. Awards: 2001-2004 Australian Postgraduate Award 2004 Cognitive Neuroscience Society Graduate Student Award 2005 Australian Psychological Society Excellent PhD Thesis Award Multiple media engagements on synaesthesia and perception Grants/Projects: Dealing with Distraction: Understanding Recovery After Interruption (2024-2028) Improving Inferences from Brain Imaging to Understand Selective Attention (ongoing) Magnetic Resonance-Compatible Audio-Visual Stimulus Delivery System (2014-) Labs/Teams: Active member of the Perception in Action Research Centre and collaborator in the Centre for Elite Performance Expertise and Training (CEPET).
Tobias Höllerer is a Professor of Computer Science at the University of California, Santa Barbara (UCSB). He leads the Imaging, Interaction, and Innovative Interfaces research group, focusing on novel user interfaces, augmented reality (AR), virtual reality (VR), and immersive visualization technologies. His work emphasizes 'Anywhere Augmentation' in AR and contributions to the Allosphere project, a three-story immersive visualization environment. Education: PhD in Computer Science, Columbia University MS in Computer Science, Columbia University Diplom (MSc equivalent), Technische Universität Berlin Research Interests: Höllerer explores AR/VR interfaces, spatial computing, human-computer interaction (HCI), and data visualization for analyzing large-scale information networks. His group develops solutions for 3D reconstruction, multimodal interaction, and applications in fields like digital fabrication, assistive technologies, and immersive storytelling. Publications Trends: Recent work spans AR navigation aids for visually impaired users, cognitive load analysis in VR, multimodal AI for spatial awareness, and creative applications like textile crafting and music instruments in AR. Scientific Awards: NSF Early Career Development Award Grants & Labs: Active in NSF-funded projects and directs the UCSB Allosphere Research Group. His lab collaborates on projects like Attention-Aware Mixed Reality Interfaces and Large-Scale Real-Time Information Visualization . Labs/Teams: Oversees the Allosphere, a unique immersive platform for data-driven exploration, and leads interdisciplinary teams advancing HCI and AR/VR technologies.
Roger Zimmermann is a Full Professor at the School of Computing, National University of Singapore (NUS), where he is also a Co-PI at the Grab-NUS AI Lab and leads the Location AI project. He previously served as Deputy Director of the NUS Smart Systems Institute (SSI) and Co-Director of the Centre of Social Media Innovations for Communities (COSMIC), both funded by Singapore’s National Research Foundation (NRF). Before joining NUS, he was a Research Area Director and Research Assistant Professor at the University of Southern California (USC). Ph.D. in Computer Science, University of Southern California (1998) M.S. in Computer Science, University of Southern California (1994) His research focuses on multimedia systems , spatio-temporal data management , streaming media architectures (especially DASH), machine learning applications , AR/VR , and location-based services . He leads the Media Management Research Lab (MMRL) at NUS, which conducts cutting-edge work in distributed multimedia and intelligent systems. His work combines theoretical depth with real-world applications in urban computing, smart mobility, and immersive media. The recent publications reflect a strong trend toward multimodal learning , spatio-temporal AI , adaptive streaming , and urban intelligence . His team explores zero-shot learning, 3D scene understanding, traffic forecasting, and open-vocabulary audio-visual segmentation, often leveraging foundational models and deep neural architectures. There is a clear emphasis on real-time, scalable systems for smart cities and immersive experiences. Dr. Zimmermann has received numerous accolades, including: DASH-IF Excellence in DASH Award (multiple years) Best Paper Awards at ACM SIGSPATIAL, IEEE ICME, and ACM MMSys Silver Award at ACM MMSys 2020 Grand Challenge IEEE Communications Society Best Editor Award (2017) ACM Distinguished Member (2017) Top 1% Publons Reviewer in Computer Science (2018) He has advised numerous students and led major research initiatives funded by MOE, NRF, A*STAR, NSF, and industry partners like Seagate, Intel, and HP. He has served as General Chair for IEEE MIPR 2023, ACM Multimedia 2020, and IEEE ISM 2015, and as TPC Co-Chair for several top-tier conferences. His editorial roles include Associate Editor for IEEE Transactions on Multimedia (TMM), ACM TOMM, and IEEE OJ-COMS. He leads the Media Management Research Lab (MMRL) , which focuses on intelligent multimedia systems, spatiotemporal data mining, and immersive media technologies. The lab develops scalable solutions for real-world challenges in urban computing, smart transportation, and interactive media.
Chuang Gan is a distinguished researcher holding dual positions as a Principal Research Staff Member at the MIT-IBM Watson AI Lab and an Assistant Professor at the University of Massachusetts Amherst. His work bridges academic research and industrial applications in artificial intelligence, with particular focus on advancing the frontiers of computer vision and multimodal learning systems. Dr. Gan's research interests span multiple interconnected domains within artificial intelligence. He specializes in video understanding, with deep expertise in representation learning, neural-symbolic visual reasoning, audio-visual scene analysis, and embodied intelligence. His work frequently integrates graph deep learning techniques with neuro-symbolic approaches to create more interpretable and robust AI systems. The recurring themes across his research portfolio include developing models that can understand physical dynamics from visual inputs, creating systems capable of embodied reasoning, and building bridges between symbolic and neural approaches to artificial intelligence. His publications reveal a strong trend toward increasingly sophisticated multimodal systems that integrate visual, auditory, and linguistic information. Over time, his work has evolved from basic video understanding tasks to complex embodied reasoning systems capable of physical simulation, 3D scene understanding, and multi-agent collaboration. A notable pattern is the progression from analyzing static scenes to understanding dynamic physical interactions and embodied agent behaviors in increasingly complex environments. Microsoft Fellowship Baidu Fellowship Dr. Gan's research has received significant recognition from major technology companies through prestigious fellowships and has been widely covered by leading media outlets including CNN, BBC, The New York Times, WIRED, Forbes, and MIT Tech Review. His work at the MIT-IBM Watson AI Lab provides him with access to substantial resources for cutting-edge AI research, while his academic position enables him to train the next generation of AI researchers. His collaborations with prominent researchers like Antonio Torralba demonstrate his integration within the top echelons of the computer vision and AI research community. At the MIT-IBM Watson AI Lab, Dr. Gan leads research initiatives focused on advancing video understanding and embodied intelligence. His work contributes to the lab's mission of developing AI systems that can perceive, reason about, and interact with the physical world in more human-like ways. His research group likely focuses on developing novel architectures for multimodal learning, creating benchmarks for physical reasoning, and building systems that can transfer knowledge between simulation and real-world environments.
Josh McDermott is a Professor in the Department of Brain and Cognitive Sciences at the Massachusetts Institute of Technology (MIT), where he also serves as Associate Department Head and previously as Interim Department Head. He leads the Laboratory for Computational Audition and conducts groundbreaking research at the intersection of psychology, neuroscience, and engineering, with a primary focus on understanding human auditory perception and its computational underpinnings. Dr. McDermott's educational background includes: PhD in Brain and Cognitive Sciences from MIT (2001-2006), advised by Edward Adelson MPhil in Computational Neuroscience from University College London (1998-2000), advised by Geoff Hinton B.A. in Brain and Cognitive Science, summa cum laude, from Harvard University (1994-1998), advised by Nancy Kanwisher McDermott's research program centers on understanding how humans derive information from sound in complex environments. His lab investigates why biological auditory systems outperform even the most sophisticated machine hearing systems in everyday situations like noisy city streets. He explores computational audition, natural sound statistics, music perception, and the relationship between sensory modalities, with long-term goals to improve treatments for hearing impairment and enable the design of machine systems that mirror human auditory abilities. His work bridges fundamental neuroscience with practical applications for hearing technologies. His recent publications reveal a strong trend toward integrating deep learning and neural network models with auditory neuroscience. His lab has been exploring how generative models, metamers, and task-optimized networks can illuminate human auditory processing. There's a clear progression from basic auditory phenomena to more complex computational models that bridge neuroscience and artificial intelligence, particularly in areas like sound segregation, music perception, and cross-modal integration. His work consistently demonstrates how computational approaches can reveal fundamental principles of auditory perception. Dr. McDermott has received numerous prestigious awards, including: Troland Research Award (2018) NSF CAREER Award (2015) APAN Young Investigator Award (2017) BCS Awards for Excellence in Undergraduate Advising (2014, 2018) BCS Award For Diversity, Equity, Inclusion and Justice (2023) As an advisor, McDermott has mentored numerous PhD students who have gone on to make significant contributions in auditory neuroscience and computational modeling. His Laboratory for Computational Audition serves as a hub for interdisciplinary research that bridges psychology, neuroscience, and engineering approaches to understanding hearing. The lab has developed innovative methodologies for studying auditory perception, including sound synthesis techniques, computational models, and cross-cultural approaches to music cognition. McDermott's work continues to push the boundaries of our understanding of auditory perception and its computational underpinnings.
Ross Greer is an Assistant Professor in the Department of Computer Science & Engineering at the University of California Merced. He leads the Mi³ Lab, focusing on machine intelligence, human-agent interaction, and safe autonomous systems. Education : B.S. and B.A. in EECS, Engineering Physics, and Music from UC Berkeley (2015), M.S. in Electrical & Computer Engineering from UC San Diego (2018), Ph.D. in Electrical & Computer Engineering at UC San Diego (2021) under Mohan Trivedi and Shlomo Dubnov. His research explores computational intelligence for open-world adaptability, robustness to rare events, and safety in chaotic environments. Key applications include autonomous driving, driver state analysis, trajectory prediction, and AI-assisted musical creativity. Recent publications emphasize vision-language models, active learning for 3D object detection, and safety metrics. Awards include the 2024 Interdisciplinary Research Award, Henry Booker Award for Ethical Engineering, and multiple best poster/grand prizes. Scientific Awards : 2024 Interdisciplinary Research Award 2024 Henry Booker Award for Exemplary Ethical Engineering Postdoctoral Networking Fellowship (Germany's Academic Exchange Service) Grand Prize (AWS Automotive Day competition at IEEE Intelligent Vehicles Symposium, 2023) Best Poster Awards (2021/2023 Jacobs Research Expo) He also co-authored the textbook Deep and Shallow: Machine Learning in Music and Audio (Chapman & Hall, 2023) and serves as a music director for UCSD's Symphonic Student Association and UC Merced's marching band.
Dr. Stavros Nousias is a researcher at the Chair of Computing in Civil and Building Engineering at the Technical University of Munich , focusing on applications of Artificial Intelligence in the Built Environment . His work bridges Knowledge Representation and Reasoning , Geometry Processing , and Machine Learning to advance construction informatics and digital twinning. Research Interests: AI for building evacuation prediction, technical drawing segmentation, BIM optimization, and respiratory disease modeling. Publications: 15+ peer-reviewed articles on topics including graph neural networks for construction simulations, pulmonary airflow analysis, and heritage site monitoring. Supervised Theses: Guided projects on AI-based BIM command prediction and robotized construction simulation . Labs: Active in the BIM-Lab and Robotic Fabrication Lab . Teaching: Co-instructor for courses like Artificial Intelligence in Engineering and Computation in Engineering 1 .
Xiaobai Liu is an Assistant Professor in the Department of Computer Science at San Diego State University, College of Sciences. His research bridges Computer Vision, Machine Learning, and Computational Statistics, with applications in Clinic Diagnosis, Sports, Transportation, Surveillance, and Video Games. Education: PhD in Computer Science from HuaZhong University of Science and Technology (2012) His research focuses on image parsing, video analysis, and deep learning techniques for 3D reconstruction, object tracking, and marine mammal detection. Recent publications highlight advancements in LiDAR generation, scene text recognition, and automated spectrogram processing. Notable grants include NSF-funded projects on autonomous vehicle safety simulations, marine mammal classification, and AI-driven recycling systems. He has advised over 30 students in real estate analytics, robotics, and computer vision projects. Scientific Awards: 2018 San Diego State University President’s Excellence Award
Shree K. Nayar is the T. C. Chang Professor of Computer Science in the School of Engineering at Columbia University, where he heads the Columbia Vision Laboratory (CAVE). He served as Department Chair from 2009-2012 and was Director of Research at Snap Inc. from 2018-2024. Nayar received his PhD from Carnegie Mellon University and has been at Columbia since 1991, progressing from Assistant to Full Professor. His educational background includes a PhD in Electrical and Computer Engineering from Carnegie Mellon University (1990), an MS from North Carolina State University (1986), and a BS from Birla Institute of Technology in India (1984). He began his career as a Research Engineer at Taylor Instruments in New Delhi before pursuing graduate studies. Nayar's research spans three interconnected areas: novel computational cameras that capture new forms of visual information, physics-based models for vision and graphics, and algorithms for scene understanding. His work in computational imaging has transformed digital photography, with applications in smartphones, robotics, virtual reality, and human-computer interfaces. His research has produced over 300 publications with nearly 60,000 citations and 80 patents. Analysis of his recent publications reveals a strong focus on computational imaging challenges including low-light vision, depth sensing, mobile interaction, and accessibility technologies. His work consistently bridges theoretical foundations with practical applications, as evidenced by commercial implementations of his assorted pixels technology in smartphone cameras. Elected to National Academy of Engineering (2008), American Academy of Arts and Sciences (2011), National Academy of Inventors (2014), and Indian National Academy of Engineering (2022) Okawa Prize (2023), IEEE PAMI Distinguished Researcher Award (2019) Two-time David Marr Prize winner (1990, 1995) - the highest honor in computer vision Multiple best paper awards at major conferences including SIGGRAPH Asia (2024) and ECCV (2024) National Young Investigator Award (1991), Packard Fellowship (1992) Nayar has supervised numerous PhD and Master's students throughout his career at Columbia. His lab has received continuous funding from NSF, industry partners, and foundations. The Columbia Vision Laboratory (CAVE) is known for its interdisciplinary approach, combining optics, hardware design, and algorithms to solve fundamental vision problems. Nayar's Bigshot Camera project demonstrates his commitment to education, providing hands-on learning experiences for students worldwide. The Columbia Vision Laboratory (CAVE) develops cutting-edge computational imaging and computer vision systems. Under Nayar's leadership, the lab has pioneered technologies including self-powered cameras, high dynamic range imaging systems, and novel computational cameras. The lab maintains strong industry connections, particularly through Nayar's role at Snap Research, and emphasizes translating research into real-world applications that benefit society.