Dr. João Henriques is a Research Fellow of the Royal Academy of Engineering (RAEng) at the Visual Geometry Group (VGG), University of Oxford. His research focuses on advancing computer vision, deep learning, and robotics, particularly in areas like 3D scene understanding, reinforcement learning, and multi-agent systems. He is renowned for developing the KCF and SiameseFC visual trackers, which won the VOT Challenge and are deployed in consumer hardware. His work spans 3D geometry, self-supervised learning, causal inference, and neuro-symbolic systems. Key contributions include methods for egocentric video analysis, unsupervised reconstruction, and robot navigation. He leads the VGG's research on neural feature fields, hierarchical scene understanding, and real-time 3D perception. Recent publications emphasize 3D-aware segmentation, universal place recognition, and neuro-symbolic world modeling for robotics. His research often bridges theoretical guarantees with practical applications, such as medical imaging and autonomous systems. Dr. Henriques collaborates with industry and academia on AI ethics, friendly AI, and interpretable learning. His lab hosts DPhil students advancing creative AI applications, such as generative models for gameplay design and LLM evaluations in real-world editorial workflows.
Martin Rajman is a Senior Scientist at École Polytechnique Fédérale de Lausanne (EPFL) with multiple affiliations across the institution. He holds positions in the School of Computer and Communication Sciences (SIN - Teaching, SCI IC MR Group, SSC - Teaching) as well as in the Vice Presidency for Strategic Development (VPS Artificial Intelligence) and the Vice Presidency for Academic Affairs (SNAI Administration). He serves as the Executive Director of Nano-tera.ch, a large Swiss Research Program funding collaborative multi-disciplinary projects in Health and the Environment. Rajman's research spans the intersection of artificial intelligence, natural language processing, and information retrieval. His work demonstrates a consistent focus on developing practical applications of computational linguistics and machine learning techniques. Early in his career, he contributed significantly to syntactic parsing, stochastic language models, and vector space representations for text. More recently, his research has expanded into deep learning applications for 3D reconstruction, empathetic conversational agents, and distributed analytics systems. His publications reveal a trajectory from foundational NLP research toward increasingly applied and interdisciplinary work connecting AI with healthcare, environmental monitoring, and human-computer interaction. Analysis of his recent publications (2015-2024) shows a clear evolution toward more applied AI research with strong interdisciplinary connections. While maintaining his core expertise in natural language processing and information retrieval, his work has expanded into computer vision, healthcare applications, and sustainable computing. The publications demonstrate increasing collaboration across disciplines, with applications in medical imaging, mental health support systems, environmental monitoring, and human-centered AI. His leadership role in the Nano-tera.ch program reflects this interdisciplinary approach, connecting computing research with real-world challenges in health and environmental contexts. Rajman has mentored several PhD students including Ailomaa Marita, Eckard Emmanuel, Melichar Miroslav, and Veselý Martin. His research has been supported through the Nano-tera.ch program, which has funded more than 100 research projects with over 95 million CHF in public funding. He has also managed more than 20 European projects during his tenure as Director of the EPFL Global Computing Center. As Executive Director of Nano-tera.ch, Rajman leads a significant research initiative connecting EPFL with national and international partners. His work bridges academic research with industry applications, notably through collaborations with eBay on product ranking technology and with Elsevier on article recommendation systems. His leadership extends to managing large-scale research programs while maintaining an active research agenda and mentoring the next generation of computer scientists.
Edward Delp is the Charles William Harrison Distinguished Professor of Electrical and Computer Engineering at Purdue University's College of Engineering. He holds affiliations with both the Department of Electrical and Computer Engineering and the Department of Biomedical Engineering. His research spans computer vision, medical imaging, and data forensics with a focus on synthetic media detection, deep learning applications, and healthcare technologies. Education: Not explicitly listed in the provided text. His work includes developing algorithms for speech forensics, microscopy image analysis, and food/nutrition assessment systems. He leads projects on synthetic speech detection, medical image segmentation, and automated crop disease measurement using RGB imaging. Delp collaborates across disciplines, integrating machine learning with healthcare and agricultural challenges. Recent work emphasizes ethical AI through fairness in synthetic media detection and explainable artifacts in biomedical imaging. He contributes to large-scale datasets like MetaFood3D and 3D nuclear segmentation frameworks for microscopy analysis. His grants and advising focus on interdisciplinary applications, though specific grant details are not provided. Delp is affiliated with the Purdue School of Biomedical Engineering and maintains active collaborations in medical imaging, computer vision, and aerospace anomaly detection.
Yang Wang is an Associate Professor in the Department of Computer Science and Software Engineering at Concordia University, holding an adjunct position since 2022. Previously, he served as an Associate Professor at the University of Manitoba (2012–2022) and worked as Chief Scientist in Computer Vision at Huawei Canada (2020–2022). He holds a PhD from Simon Fraser University, MSc from the University of Alberta, and BEng from Harbin Institute of Technology. His research focuses on computer vision, machine learning, and deep learning, particularly in meta-learning, test-time training, and continual learning. Key areas include crowd counting, anomaly detection, video highlight detection, and gaze estimation. His work has been recognized with awards such as the Falconer Emerging Researcher Rh Award (2017) and a Faculty of Science Research Chair (2019–2022). Recent research emphasizes AI models that are personalized and adaptable, leveraging techniques like meta-learning and few-shot learning. He has published extensively in top venues (CVPR, ICCV, ECCV) and holds patents in related fields. His group collaborates with industry partners like Huawei and Sightline Innovation.
Yuan-Fang Li is an Associate Professor in the Department of Data Science & AI at Monash University's Faculty of Information Technology. He also serves as Associate Dean International. His research focuses on knowledge graphs, natural language processing, multimodality, and graph representation learning. He holds a PhD from National University of Singapore (2006) and a Bachelor of Computing (Honours) from the same institution (2002). Affiliations: Monash University (since 201?), National University of Singapore (PhD 2002-2006) Key Projects: Leading research on neuro-symbolic systems (HARNESS project), large-scale multimodal knowledge management, and maritime knowledge graphs Teaching: Taught courses including FIT4002, FIT4004, and supervised over 20 PhD students Research interests include complex question answering over knowledge graphs, knowledge extraction from text/images, and structural/temporal graph learning. He has published 152+ works with notable contributions to scene graph generation, event extraction, and LLM-based reasoning. Key awards include the 2020 Best Student Paper Award and 2017 Kurzweil Prize. Grants: ARC Discovery Projects, industry collaborations (e.g., Outotec Oy) Labs/Teams: Active in Monash's Data Science & AI research groups, leading neuro-symbolic AI initiatives
Haoyi Xiong is an active academic researcher in artificial intelligence, machine learning, and data science, with extensive publications in top-tier journals and conferences including IEEE TPAMI, NeurIPS, ICML, KDD, and AAAI. His work spans explainable AI, graph neural networks, diffusion models, remote sensing, and large language models. Research Interests: Explainable AI (XAI) and model interpretability Graph Neural Networks and contrastive learning Diffusion models and generative AI Medical and remote sensing image analysis Large language models and autonomous agents Learning to rank and web search His recent publications (2023–2025) show a strong trend toward self-supervised learning , model robustness , and integration of LLMs with structured data and knowledge graphs . He frequently collaborates with researchers from major tech and academic institutions. Scientific Awards: No explicit awards mentioned in the provided text. Advising and Grants: While no direct mention of students or grants, his role as a senior author on numerous papers suggests he advises graduate students and likely leads funded research projects in machine learning and AI. His work on frameworks like COLTR , GS2P , and MUSCLE indicates leadership in developing scalable AI systems. Labs and Teams: Though not explicitly stated, his frequent collaboration with Jiang Bian, Dejing Dou, and Dawei Yin suggests affiliation with a well-established AI research lab or industry-academia partnership focused on data mining, intelligent systems, and large-scale learning.
Yihao Ding is a Research Fellow at the School of Physics, Mathematics and Computing at The University of Western Australia. He holds a Ph.D. in Computer Science from the University of Sydney, awarded on November 11, 2024. Dr. Ding's educational background includes: Doctor of Philosophy in Computer Science, Visually Rich Document Understanding and Intelligence, University of Sydney (March 1, 2021 - November 11, 2024) His research focuses on multimodal large language models, deep learning-based document analysis, information retrieval, question answering, and interdisciplinary applications of deep learning. Dr. Ding has published extensively in leading conferences and journals, including ACL, CVPR, AAAI, IJCAI, SIGIR, ECML-PKDD, COLING, and CIKM. His current work spans visual document understanding, multimodal learning, natural language processing, and interdisciplinary applications including geographic information systems. Dr. Ding's recent publications demonstrate a strong focus on visually-rich document understanding, multimodal learning, and interdisciplinary applications. His work ranges from developing novel multimodal models for form document understanding to creating comprehensive datasets for visual question answering and applying machine learning to environmental challenges like lithium recovery from water sources. His research shows a consistent pattern of addressing complex multimodal problems with innovative deep learning approaches. Dr. Ding is an active member of the AI community, having organized workshops, tutorials, and competitions at top-tier venues such as AAAI, IJCAI, and CIKM. He has also served as a Chair or Reviewer for major conferences including IJCAI, ARR Rolling, ICLR, ACMMM, CVPR, ICCV, WACV and IJCNN.
Rynson W.H. Lau is a Professor of Computer Science at City University of Hong Kong (CityU), leading research in Computer Graphics, Computer Vision, and Deep Learning. He holds an Honorary Professorship at Swansea University. Previously, he served on faculties at Durham University and The Hong Kong Polytechnic University. His work focuses on advancing graphics and vision techniques, including deep learning applications for graphics/vision problems, with publications in top venues like SIGGRAPH, CVPR, and NeurIPS. He has received the Adobe Research Gift (2023) and the Springer Nature Editorial Contribution Award (2025) for his editorial contributions to the International Journal of Computer Vision . Education: B.Sc. (First-class Honors) in Computer Systems Engineering from University of Kent Ph.D. in Computer Science from University of Cambridge Research Interests: Computer Graphics: Focused on 3D reconstruction, rendering, and real-time performance capture. Computer Vision: Specializing in saliency detection, object recognition, and low-light scene enhancement. Deep Learning: Developing generative models and diffusion-based frameworks for graphics and vision tasks. Editorial Roles: Editorial Board Member, International Journal of Computer Vision and IET Computer Vision . Guest Editor for special issues in journals like ACM Transactions on Internet Technology and IEEE Transactions on Multimedia. Teaching: 2024/25 Academic Year: CS4185: Multimedia Technologies and Applications CS4188/CS5188: Virtual Reality Technologies and Applications Research Team: Advises over 20+ students and collaborates internationally. Recent projects include AI-driven VR systems for healthcare and advanced 3D content generation using diffusion models.
Chris Thomas is an Assistant Professor in the Department of Computer Science at Virginia Tech’s College of Engineering. His research focuses on computer vision, cross-modal retrieval, and multimodal knowledge representation, with applications in information extraction, fake news detection, and AI safety. He leads the Sanghani Center for Artificial Intelligence and Data Analytics and has received grants from the Commonwealth Cyber Initiative and a Google Research Scholar award. Education: Ph.D. (2020) and B.S. (2013) in Computer Science, University of Pittsburgh. Research Interests: Developing robust cross-modal systems that bridge vision and language. Key areas include fine-grained visual entailment, multimodal inconsistency detection, and defending AI models against adversarial attacks. His work emphasizes practical applications like fact-checking and cybersecurity. Recent Contributions: Led the development of JourneyBench (a vision-language benchmark) and the Semantic Shield defense framework. Recent grants include cybersecurity for embodied agents and safer multimodal web agents. Awards: Google Research Scholar Award (2025), multiple Commonwealth Cyber Initiative grants (2024–2025). Advising & Grants: Advises students like Hani Alomari (ACL 2025) and collaborates with institutions like Columbia University and UCLA. His work integrates multimodal data to address real-world challenges in AI safety and information integrity. Labs/Teams: Active in Virginia Tech’s AI initiatives, focusing on interdisciplinary research at the intersection of vision, language, and security.
Fabrizio Falchi is a researcher at the Artificial Intelligence for Media and Humanities (AIMH) Lab of the Institute of Information Science and Technologies (ISTI) within Italy's National Research Council (CNR). He also maintains an associate position at the Biorobotics Institute of Scuola Superiore Sant'Anna. His work focuses on developing advanced multimedia retrieval systems, with the VISIONE platform being his most notable contribution, which has won international competitions including the Video Browser Showdown in 2024 and placed second in 2023. Falchi's educational background includes: Ph.D. in Information Engineering from University of Pisa (Italy) Ph.D. in Informatics from Faculty of Informatics of Masaryk University of Brno (Czech Republic) M.B.A. from Scuola Superiore Sant'Anna in Pisa His research spans deep learning, convolutional neural networks, deep features extraction, similarity search algorithms, distributed indexing systems, multimedia information retrieval, computer vision applications, and peer-to-peer systems. Falchi has made significant contributions to fine-grained visual understanding, cross-modal retrieval (particularly image-text matching), and robustness of deep learning systems against adversarial attacks. His work demonstrates a strong focus on practical applications of these technologies, particularly in video retrieval systems and safety monitoring solutions. Analysis of Falchi's recent publications reveals a strong focus on video and image retrieval systems, with the VISIONE platform being central to his work. His research shows increasing emphasis on fine-grained understanding in computer vision, cross-modal retrieval, and addressing practical challenges like cross-resolution face recognition. Recent work demonstrates innovation in making these systems more efficient through techniques like knowledge distillation (ALADIN) and leveraging virtual worlds for training data. His publications consistently bridge theoretical advances with practical applications in surveillance, safety monitoring, and multimedia search. Falchi's work has received significant recognition: Best paper award at CBMI 2024 for 'Is ClLIP the main roadblock for fine-grained open-world perception?' VISIONE 2024 won the Video Browser Showdown competition in Amsterdam VISIONE obtained second place at Video Browser Showdown 2023 in Bergen Best Paper Award for 'Learning Safety Equipment Detection using Virtual Worlds' at CBMI 2019 Falchi collaborates extensively with researchers at ISTI-CNR, particularly within the AIMH Lab. His work on VISIONE involves collaboration with Giuseppe Amato, Paolo Bolettieri, Fabio Carrara, Claudio Gennaro, Nicola Messina, Lucia Vadicamo, and Claudio Vairo. As co-chair of Ital-IA 2023, the 3rd National Conference on Artificial Intelligence, he plays an active role in the academic community. He is a member of ACM (since 2012), the Computer Vision Foundation, the Italian Association for Computer Vision Pattern Recognition and Machine Learning (CVPL), and the CINI Lab on Artificial Intelligence and Intelligent Systems. Falchi is a key member of the Artificial Intelligence for Media and Humanities (AIMH) Lab at ISTI-CNR, where he leads research on video retrieval systems. The lab has developed the award-winning VISIONE platform, which combines multiple scientific results in content-based video retrieval. His team focuses on developing systems that enable users to search for target videos using textual prompts, drawing objects and colors, or images as query examples. The lab's work demonstrates strong interdisciplinary collaboration, bridging computer science with practical applications in media, safety monitoring, and urban environments.
Dr. Armin Mustafa is an Associate Professor in Computer Vision and AI at the University of Surrey, where he holds a prestigious Royal Academy of Engineering Research Fellow position. He is affiliated with the Centre for Vision, Speech and Signal Processing (CVSSP), the School of Computer Science and Electronic Engineering, and the Surrey Institute for People-Centred Artificial Intelligence (PAI). His research focuses on developing AI systems for visual understanding of complex dynamic scenes, with applications in entertainment, autonomous systems, and augmented/virtual reality. Dr. Mustafa completed his PhD in general dynamic scene reconstruction from multi-view videos in 2016 from the University of Surrey under the supervision of Prof. Adrian Hilton. Prior to his doctoral studies, he worked for three years (2010-2013) at Samsung Research Institute in Bangalore, India, in the field of Computer Vision. His research expertise spans Computer Vision, Scene Understanding, 3D/4D Vision, Virtual Reality, Light Fields, Machine Learning, Video Captioning, Augmented Reality, Artificial Intelligence, and Audio-visual Video Understanding. Dr. Mustafa has pioneered advances in 4D vision, NLP, and Scene Understanding over the past decade, with a particular focus on enabling machines to model and interpret real-world environments for socially beneficial applications. His work bridges theoretical advances in computer vision with practical applications in media production, virtual reality, and autonomous systems. Analysis of Dr. Mustafa's recent publications reveals a strong focus on multimodal learning, particularly the integration of audio and visual information for scene understanding. His work spans diverse areas including shadow detection and removal, audio event classification, video captioning, person image generation, and dynamic scene reconstruction. A notable trend is his exploration of transformer architectures for both vision and audio tasks, as well as the application of self-supervised learning techniques to reduce dependency on labeled data. Dr. Mustafa has received numerous prestigious awards: 2018 - Research Fellowship, The Royal Academy of Engineering, UK 2017 - Young Researcher award, CVPR 2016 - Doctoral Consortium grant, CVPR 2015 - BMVA travel grant for ICCV 2014 - Set-Squared Research to Innovator grant 2013 - Overseas Research Scholarship, FEPS, The University of Surrey 2010 - Cadence Silver Medal, Indian Institute of Technology, Kanpur As a dedicated mentor, Dr. Mustafa supervises several PhD students working on cutting-edge topics including multi-person reconstruction, audio-visual scene understanding, and automatic storyboard generation. His research is supported by significant grants including a £15 million UKRI Prosperity Partnership with the BBC (AI4ME), a 5-year Royal Academy of Engineering fellowship (4D Vision for Perceptive Machines), and multiple projects with industry partners such as Figment Productions and Foundry. Dr. Mustafa is an active member of the Centre for Vision, Speech and Signal Processing (CVSSP), one of the world's leading research centers in vision, speech, and signal processing. He also contributes to the Surrey Institute for People-Centred Artificial Intelligence (PAI), where he serves as a Surrey AI Fellow. His work often involves collaboration with industry partners and other academic institutions across Europe.
Tetsunori Kobayashi is a Professor in the School of Fundamental Science and Engineering at Waseda University, Japan, where he has served since 1997. He is renowned for pioneering research in human–robot interaction, spoken language processing, and multimodal conversational systems, leading to over 230 refereed papers and an h-index of 35 (Google Scholar). Education: 1980 B.Eng. in Electrical Engineering, Waseda University 1982 M.Eng. and 1985 Dr.Eng. from Graduate School of Science and Engineering, Waseda University Research Interests: His work spans intelligent robotics , perceptual information processing , pattern recognition , image and audio processing , and conversational AI . He develops algorithms for real-time dialogue systems, multi-party conversation facilitation robots, and non-autoregressive speech recognition leveraging CTC and pre-trained language models. Recent Publication Trends: Since 2020 his group has advanced non-autoregressive end-to-end ASR (Mask-CTC, Intermpl, BECTRA), noise-robust attention , multi-look-ahead conversational ASR , and neural speaker diarization . They integrate BERT-style pre-training with CTC losses to accelerate inference while maintaining accuracy. Parallel work explores vision-and-language topics such as scene-graph generation, video semantic indexing, and personalized summarization for spoken news delivery. Scientific Awards: IEICE Fellow 2023 – for multi-modal multi-party conversation research IPSJ Fellow 2016 – for pioneering robot conversation studies JST Award for Academic Start-ups 2024 Best Paper Awards from IEICE, IEEE BTAS, ACM SIGGRAPH VRCAI, and several IPSJ workshop prizes Advising & Grants: He has mentored dozens of PhD and Master’s students who now lead in academia and industry. Major funded projects include JST CREST on conversational robotics, NEDO and JST-support for AI-based speech interfaces, and industry collaborations with NHK, OKI, and NEC. Labs & Teams: Kobayashi heads the Perceptual Computing Laboratory at Waseda, conducting interdisciplinary research with domestic and international partners such as MIT, ATR, and NHK Science & Technology Labs.
Dr. Xiaohan Yu is a Lecturer in Artificial Intelligence at Macquarie University's School of Computing, joining in December 2023. Previously, he completed his doctoral studies at Griffith University and served as a Research Fellow at the ARC Research Hub for Driving Farming Productivity. His research focuses on Ultra-Fine-Grained Visual Categorization (Ultra-FGVC), Smart Farming, and Automated Crop Cultivar Identification, with over 70 publications in top-tier venues like ICCV, CVPR, and IEEE Transactions. He holds editorial roles at Pattern Recognition and SN Computer Science , and received the APRS Early Career Award (2022) and ACM MM 2024 Outstanding Area Chair distinction. Education: Completed doctoral studies in Artificial Intelligence at Griffith University, Australia. Research Interests: Ultra-Fine-Grained Visual Categorization (Ultra-FGVC) Smart Farming and Agricultural Robotics Computer Vision Applications in Healthcare (e.g., trachoma detection) Deep Learning, Continual Learning, and Domain Adaptation Key Contributions: Pioneered Ultra-FGVC research, developed frameworks like Mix-ViT and CLE-ViT, and contributed to benchmarking multi-object tracking in farming. His work bridges pattern recognition with real-world applications in agriculture and healthcare. Scientific Awards: Australian Pattern Recognition Society (APRS) Early Career Researcher Award 2022 ACM Multimedia 2024 Outstanding Area Chair Award Advising & Grants: Actively involved in editorial roles (Area Chair for ACM MM, IJCNN) and grant-funded research through ARC hubs. His work is supported by collaborations in agriculture and AI-driven solutions for crop cultivar identification. Labs & Affiliations: Member of Macquarie's Smart Green Cities Research Centre and Frontier AI Research Centre , advancing interdisciplinary AI applications.
Archontis Politis is an Assistant Professor in the Department of Computing Sciences at Tampere University's Faculty of Information Technology and Communication Sciences. His research focuses on signal processing, machine learning, and their applications in audio engineering, particularly in spatial audio, sound source separation, and parametric audio coding. He explores topics such as Ambisonics, reverberation control, and neural network-based approaches for audio processing. His work emphasizes spatial audio reproduction, including six degrees of freedom (6DOF) rendering, microphone array processing, and efficient compression techniques for higher-order Ambisonics. He also investigates sound event localization and detection, leveraging machine learning for real-world acoustic scenarios. His contributions span theoretical advancements in spherical harmonics and practical implementations of spatial audio systems. Recent research highlights include developing datasets for music source separation, improving synthetic-to-real generalization in classical music, and creating neural encoding models for irregular microphone arrays. His methodologies often integrate deep learning with traditional signal processing to address challenges in multi-speaker environments and dynamic acoustic scenes.
Professor Matt Garratt is a faculty member at the University of New South Wales (UNSW Canberra), School of Engineering and IT, serving as AI theme lead for the Defence Trailblazer Universities initiative with over $200 million in funding. His primary research focuses on sensing, guidance, and control for autonomous systems within robotics and unmanned aerial vehicles. Garratt's research spans robotics, swarm intelligence, and autonomous systems with emphasis on bio-inspired navigation techniques and adaptive flight control. His work addresses critical challenges including terrain following using vision systems, landing UAVs on moving platforms, and developing self-organizing swarms. He integrates artificial intelligence, computer vision, and machine learning to advance unmanned systems capabilities in complex environments. Analysis of his recent publications reveals strong trends in bio-inspired UAV navigation (particularly honeybee behavior modeling) and swarm robotics applications. His work increasingly incorporates deep learning for perception tasks while addressing real-world challenges like gas plume detection and adversarial robustness in 3D vision systems. The research demonstrates consistent progression toward practical implementation of autonomous systems in dynamic environments. Professor Garratt has secured over $7.7 million in external research funding as Chief Investigator on 33 grants. He actively mentors graduate students with scholarships available for Masters and PhD research in robotics and AI, focusing on: UAV path planning and adaptive control systems Swarm robotics collective motion optimization Bio-inspired autonomous navigation techniques Computer vision for robotic perception He co-founded the UNSW Canberra AIR (AI and Robotics) Group (AIR Lab), which drives research in trusted autonomy, swarm intelligence, and AI integration for defense applications. The lab develops practical solutions for autonomous systems operating in complex, real-world environments while maintaining ethical AI frameworks.