Christos Faloutsos is the Fredkin Professor of Computer Science at Carnegie Mellon University, with a courtesy appointment in Electrical and Computer Engineering. He holds a B.Sc. from the National Technical University of Athens and M.Sc./Ph.D. from the University of Toronto. His research focuses on data mining, graph analysis, fractals, and database systems. Notable contributions include foundational work on R-trees, graph mining laws (e.g., Kronecker graphs), and applications in medical imaging, network security, and fraud detection. Key projects include PEGASUS (petascale graph mining), fraud detection in online auctions (NetProbe), and tools for human trafficking analysis (TrafficVis). He has led NSF-funded projects on tensor mining, network anomaly detection, and bioinformatics. Over 300 refereed publications highlight his contributions across databases, data mining, and networks. Awards include the KDD Best Paper (2005, 2016), SIGMOD Test-of-Time Award, and recognition as a top nurturer in IT. His lab collaborations span the Parallel Data Lab (PDL), Machine Learning Department, and Computational Biology.
Andy Pavlo is an Associate Professor with Indefinite Tenure in the Computer Science Department at Carnegie Mellon University's School of Computer Science. He is an active member of the CMU Database Group and the Parallel Data Laboratory, where he leads research in database management systems with a focus on self-driving architectures, transaction processing, and large-scale analytics. His work bridges academic research and industry applications through projects like NoisePage, OtterTune (which he co-founded and served as CEO before it ceased operations), and Peloton. Dr. Pavlo's research interests span database management systems with particular emphasis on autonomous database architectures that can self-tune and optimize without human intervention. His work explores transaction processing systems that can handle high-throughput workloads while maintaining consistency, and large-scale data analytics techniques that efficiently process massive datasets. He has made significant contributions to query optimization, database extensibility, and automatic database tuning using machine learning techniques. His recent work on database extensibility revealed critical issues in PostgreSQL's extension ecosystem, showing that approximately 16% of extensions are incompatible with at least one other extension due to API violations and memory errors. His research output demonstrates a consistent focus on practical database systems challenges, with recent publications examining database extensibility, user-defined function optimization, and the cyclical nature of database research. The articles show a strong trend toward making database systems more autonomous, with increasing integration of machine learning techniques for automatic tuning and optimization. His work often combines deep theoretical analysis with practical implementation in open-source systems. Dijkstra Award 2024 for contributions to database systems research Dr. Pavlo actively mentors graduate students, with current advisees including Wan Shen Lim, William Zhang, and Sam Arch (co-advised with Todd Mowry). His former students have gone on to successful careers in both industry and academia. He has secured significant research funding through CMU's affiliate program with major database companies including ClickHouse, DataStax, dbt, Firebolt, MotherDuck, RelationalAI, SingleStore, Spiral, PingCAP/TiDB, Yellowbrick, and Yugabyte. His research is supported by these industry partnerships and likely includes NSF funding given his active participation in the database research community. At CMU, Dr. Pavlo leads the Database Group and organizes several seminar series including "SQL or Death," "Database Building Blocks," and "ML⇄DB Technical Talks." These seminars bring together researchers and practitioners to discuss cutting-edge developments in database systems. He also runs a summer research internship program that has attracted students for multiple consecutive years, indicating a strong research group with ongoing projects and funding.
Joy Arulraj is an Associate Professor in the School of Computer Science within the College of Computing at Georgia Institute of Technology. His research focuses on data systems, machine learning, and database systems, with a particular emphasis on video analytics and adaptive query processing. He leads the Data Systems and Analytics Group and is developing the EVA AI-Relational Data System. Dr. Arulraj's research interests span data systems, machine learning, database systems, video analytics, and adaptive query processing. His work centers on developing systems that efficiently process complex queries, particularly for video analytics and machine learning workloads. He has made significant contributions to GPU database systems, non-volatile memory database management, and adaptive query processing techniques. His research often bridges the gap between theoretical database principles and practical implementations for modern hardware architectures. His recent publications show a strong trend toward video analytics systems, adaptive query processing for machine learning workloads, and GPU-accelerated database systems. The EVA system represents a major focus of his recent work, providing end-to-end exploratory video analytics capabilities. His research also addresses fundamental database concepts like buffer management, query optimization, and storage management, adapting these principles for modern hardware and application requirements. Dr. Arulraj has advised numerous graduate students including Pramod Chunduri, Gaurav Tarkok Kakkar, Jiashen Cao, and Sayan Sinha. His graduated students have gone on to work at companies like ServiceNow, Meta Research, and the Korean Army. He actively teaches database system courses at Georgia Tech, including Database System Implementation (CS 4420/6422) and Advanced Database System Implementation (CS 4423/6423), where students build database systems from scratch using C++ and the BuzzDB framework. He maintains an active research program with consistent publication output across top database and systems conferences. His work spans from theoretical database principles to practical system implementations, with a recent emphasis on video analytics, machine learning integration with database systems, and leveraging modern hardware like GPUs and non-volatile memory for database applications.
Lin Ma is currently an Assistant Professor at the University of Michigan, Ann Arbor in the Department of Electrical Engineering and Computer Science (College of Engineering). Previously, they served as a Post Doctoral Fellow at Carnegie Mellon University (2021-2022) and as a Software Engineer at Databricks, Inc. (2022-2023). Research Interests focus on the intersection of database systems and machine learning, particularly in developing self-driving database management systems . Key areas include workload forecasting , automated index optimization , query execution acceleration , and machine learning integration for database automation. Their work explores GPU-accelerated analytics, memory optimization, and transactional consistency models. Academic Contributions span 15+ publications in top venues like VKDB , SIGMOD , and CIDR , including recent 2025 papers on Vortex (GPU memory optimization) and Scompression (workload compression). Earlier work introduced QueryBot 5000 , a workload forecasting framework, and explored anti-caching for storage optimization in OLTP systems. Teaching includes courses like EECS 584: Advanced Database Management Systems and EECS 484: Database Management Systems at the University of Michigan (2023-2025), and 15-445/645 Database Systems at Carnegie Mellon University. Service involves program committee roles for SIGMOD (2023-2025), VLDB (2022-2025), and CIDR (2024-2025). They also served on admissions and search committees at both institutions. Advising includes supervising PhD and MS students: Siyuan (Doug) Dong , Zhongwei Xu , and Haotian (Jack) Gong (co-advised with Barzan Mozafari), among others.
Brendan T. O'Connor is an Associate Professor at the College of Information and Computer Sciences, University of Massachusetts Amherst, where he directs the SLANG Lab and serves as Associate Director of the Computational Social Science Institute. His research bridges statistical machine learning and natural language processing with social science applications, particularly using text data from news and social media to understand societal patterns. His work focuses on developing text analysis methods to answer social science questions in domains like political science and sociolinguistics. Current collaborative projects include combating misinformation, analyzing bias in news coverage, and developing tools for clinical discourse assessment using large language models. O'Connor's publications demonstrate a consistent focus on computational social science, with recent work exploring multilingual analysis, legal discourse patterns, and sociolinguistic variation. His methodological contributions span coreference resolution, event extraction, and argument mining. At UMass, he contributes to multiple research centers including the Computational Social Science Institute, UMass NLP group, and Centers for Data Science and Intelligent Information Retrieval. He teaches graduate seminars in natural language processing and maintains active collaborations across disciplines.
Phillip B. Gibbons is a Professor in both the Computer Science Department and Electrical & Computer Engineering Department at Carnegie Mellon University. He received his Ph.D. in Computer Science from the University of California at Berkeley in 1989 and has held research positions at AT&T Bell Laboratories, Lucent Bell Laboratories, and Intel Research Pittsburgh before joining CMU's faculty. His research spans parallel computing, distributed systems, databases, computer architecture, and machine learning. Gibbons' work bridges theory and systems, with publications in top-tier conferences including SOSP, OSDI, SIGMOD, VLDB, NeurIPS, and many others across computer science and engineering disciplines. His research has been supported by significant funding from NSF, Intel, and other organizations. Gibbons has made substantial contributions to streaming algorithms, parallel computing frameworks, distributed systems security, and large-scale machine learning systems. His work on data stream algorithms with Alon, Matias, and Szegedy has been particularly influential in the field. He has served in numerous leadership roles including Editor-in-Chief of ACM Transactions on Parallel Computing (2012-2018) and on the editorial boards of Journal of the ACM and IEEE Transactions on Cloud Computing. He has also been active on program committees for major conferences in systems, databases, and theory. IEEE Fellow (2014) - For contributions to parallel computing and databases ACM Fellow (2006) - For contributions to parallel computing, databases, and sensor networks Selected for Oral Presentation at NeurIPS '13 (only 20 selected out of 1420 submissions) Co-winner of the best paper award for NSDI '06 Gibbons has advised numerous students and mentored researchers who have gone on to make significant contributions in academia and industry. His research has been supported by major grants including the $15M Intel Science and Technology Center for Cloud Computing (2011-2015) where he served as Co-PI/Co-Director. He currently leads research projects on write-efficient algorithms, big learning systems, and visual cloud systems. His laboratory work focuses on bridging theoretical computer science with practical systems implementation, particularly in the areas of parallel and distributed computing. Current research directions include adapting algorithms for emerging memory technologies and optimizing machine learning systems for large-scale deployment.
Carlos Guestrin is the Fortinet Founders Professor of Computer Science at Stanford University and serves as Director of the Stanford AI Lab (SAIL) and Senior Fellow at the Institute for Human-Centered AI (HAI). He holds dual roles as Chief Scientist at Visual Layer and Virtue AI. His research focuses on machine learning methods, explainability, fairness, and ethics of AI, alongside systems for scalable AI deployment. Education details are not explicitly provided, but his work spans foundational contributions to machine learning systems (e.g., XGBoost) and explainable AI frameworks like Anchors and LIME. He emphasizes ethical AI through projects like CheckList for model testing and Model Equality Testing for API transparency. His scientific contributions include advancing optimization techniques (AdaScale SGD, TVM compiler) and ethical benchmarks for generative AI. He has been recognized as a Member of the National Academy of Engineering for his transformative impact on AI systems and their societal applications. Guestrin leads interdisciplinary initiatives at SAIL and HAI, fostering collaboration between technical innovation and human-centered design. His work bridges theory and practice, addressing challenges in healthcare (diabetes management systems) and AI security.
Joel Greenhouse is a Professor of Statistics at Carnegie Mellon University (CMU), affiliated with the Department of Statistics & Data Science. He has been on the faculty since 1983 and held leadership roles, including serving as Associate Dean of the College of Humanities and Social Sciences from 1997 to 2002. He also holds an adjunct appointment as Professor of Epidemiology and Psychiatry at the University of Pittsburgh. His expertise spans statistical methodology, clinical trial design, and meta-analysis, with a focus on integrating data from multiple sources to address complex healthcare and public health challenges. Greenhouse earned his Ph.D. in Biostatistics from the University of Michigan and completed a postdoctoral fellowship at CMU. His research emphasizes developing statistical tools for observational studies, clinical trials, and meta-analytic frameworks, particularly in neurology, mental health, and public policy contexts. Notable contributions include analyzing the impact of media on youth suicide rates, improving aphasia classification through automated speech analysis, and evaluating highway safety through driver health data. Education: Ph.D. in Biostatistics, University of Michigan Affiliations: Adjunct Professor at University of Pittsburgh, Member of National Academy of Sciences’ committees Professional Service: Data and safety monitoring boards for NIH/VA studies, co-chair of Federal Motor Carrier Safety Administration review panels His awards include CMU’s Doherty Award for Education, Ryan Teaching Award, and E. Dunlop Smith Award for teaching excellence. His work bridges theoretical statistics with real-world applications, particularly in interdisciplinary collaborations across medicine, psychology, and public policy. Greenhouse’s recent articles highlight trends in leveraging large datasets for clinical insights (e.g., aphasiaBank), re-evaluating environmental and behavioral health associations, and advancing causal inference methods. His interdisciplinary approach ensures statistical rigor addresses societal challenges, from suicide prevention to highway safety.
S. Mohadeseh Taheri-Mousavi is an Assistant Professor in the Department of Materials Science and Engineering at Carnegie Mellon University (CMU), part of the College of Engineering. She joined CMU in September 2022 after postdoctoral appointments at MIT and Brown University. Her research is supported by major grants from NASA STRI, DARPA, the Army Research Laboratory, and the Naval Nuclear Laboratory, and she is affiliated with the NextManufacturing Center and the Wilton E. Scott Institute for Energy Innovation. Her educational background includes a Ph.D. from EPFL, Switzerland, and M.Sc. and B.Sc. degrees from Sharif University of Technology, Iran. She was awarded both early and advanced Swiss National Science Foundation fellowships during her postdoctoral studies. Taheri-Mousavi’s research focuses on the intersection of materials science, mechanical engineering, and computer science. She develops multi-scale computational models and AI-driven frameworks—such as AlloyGPT and generative AI agents—to design next-generation structural alloys, particularly for additive manufacturing and extreme environments. Her work emphasizes materials sustainability, industrial decarbonization, and uncertainty quantification in alloy design. The integration of machine learning with Integrated Computational Materials Engineering (ICME) and CALPHAD methods enables rapid exploration of high-dimensional composition and processing spaces. Her recent publications (2023–2025) show a strong trend toward AI/ML applications in alloy discovery, hydrogen embrittlement modeling, and high-temperature aluminum and tungsten alloys. These works reflect a deep commitment to accelerating materials innovation through human-AI collaboration and smart experimental validation. Her scientific honors include prestigious Swiss National Science Foundation fellowships. She has also received seed funding from the Scott Institute for Energy Innovation to study hydrogen embrittlement. She advises a dynamic team of doctoral students and a postdoctoral researcher, working on topics including hydrogen embrittlement, generative AI for welding, and gradient alloys. Her research is funded by high-impact grants from NASA, DARPA, the Army, and the Naval Nuclear Laboratory, supporting transformative projects in structural alloy design. She leads the Taheri-Mousavi Group, which operates within CMU’s Materials Characterization Facility and the NextManufacturing Center. The group focuses on developing novel AI-integrated computational frameworks to guide efficient and intelligent experimentation in alloy development.
Rebecca Nugent is the Stephen E. and Joyce Fienberg Professor of Statistics & Data Science and Department Head at Carnegie Mellon University. She holds a PhD in Statistics from the University of Washington (2006), an MS in Statistics from Stanford (2006), and a BA in Mathematics, Statistics, and Spanish from Rice University (2002). Her research spans clustering methodology , record linkage , educational data mining , public health , and semantic organization , with a focus on high-dimensional data and adaptive learning environments. She leads the Integrated Statistics Learning Environment (ISLE) and Corporate Capstone programs, emphasizing low-barrier data platforms for education and industry collaboration. Academic Roles : Department Head, Carnegie Mellon; Affiliated Faculty, Block Center for Technology and Society Research Grants : NSF (2017-2019), NIH (2018), Carnegie Mellon ProSEED/Simon Initiative (2020, 2018), Berkman Fund (2014) Her 15 most recent publications focus on data science pedagogy, clustering algorithms, record linkage applications in historical and medical data, educational data mining, and semantic organization studies. Awards include the ASA Waller Education Award (2015) and the William H. and Frances S. Ryan Award (2015) . She mentors a diverse group of PhD, Master's, and undergraduate students, with alumni pursuing careers in academia, industry, and sports analytics.
Yihan Sun is an Assistant Professor at the University of California, Riverside (UCR) since January 2020. He earned his Ph.D. in Computer Science from Carnegie Mellon University (CMU) , advised by Guy Blelloch , and holds a Bachelor's degree in Computer Science from Tsinghua University . Research Interests: Yihan Sun focuses on the theory and practice of parallel computing , including Parallel algorithms and data structures Write-efficient algorithms for Non-Volatile Memory (NVM) Computational geometry (range trees, Delaunay triangulations) Graph algorithms (SSSP, SCC, cluster-based BFS) Concurrent and persistent data structures Multi-version concurrency control (MVCC) with garbage collection Applications in databases, transactional systems, and computational biology Recent Research Trends: His work on join-based parallel balanced trees has been foundational, supporting four balancing schemes (AVL, red-black, weight-balanced, treaps) and enabling efficient implementations in graph analytics, spatial queries, and dynamic programming. Recent publications focus on output-sensitive algorithms , scalable graph libraries (PASGAL) , and pedagogical approaches to teaching parallel algorithms. Teaching: He teaches CS260 (Parallel Algorithms) at UCR and has served as a guest lecturer for MIT 6.886 (Algorithm Engineering) and CMU 15-859 (Algorithms in the real world) . He also contributed to algorithm education through a tutorial at the ACM Symposium on Principles and Practice of Parallel Programming (PPoPP 2019) . Labs & Collaborations: Yihan is a core contributor to the PAM (Parallel Augmented Maps) library, which has been integrated into systems like Aspen (graph-streaming) and C-trees . He collaborates with teams at CMU-Parlay , PBBS , and Ligra , with his code available on Github for community feedback.
Andrew O. Arnold is a Principal Applied Machine Learning Engineer at Shopify and an Adjunct Professor at New York University's Tandon School of Engineering, Department of Finance and Risk Engineering. He earned his Ph.D. in Machine Learning from Carnegie Mellon University and a BA in Computer Science and Artificial Intelligence from Columbia University. Education Ph.D., Machine Learning, Carnegie Mellon University BA, Computer Science and Artificial Intelligence, Columbia University His research focuses on robust machine learning , developing models that perform well in low signal-to-noise regimes, handle distributional shifts (transfer learning), and extract features from unstructured data. Key applications include time series analysis and natural language processing in financial and other domains. Recent publications highlight work on large language models (LLMs) for code generation, including multitask pretraining, contrastive learning, and quantization techniques for efficiency. He has contributed to understanding model robustness and adapting NLP methods to dynamic market conditions. Arnold teaches NYU FRE GY 7871: News Analytics and Machine Learning , covering NLP and ML techniques for quantitative trading strategies. The course emphasizes practical applications of sentiment analysis, text relevance, and novelty detection in financial contexts. He has led teams at Amazon Web Services (AI Labs), served as Chief Scientist at Oracle Alpha, and worked at Microsoft Research, IBM Research, and other institutions. His technical expertise spans code generation , anomaly detection , and NLP for commerce , with patents in these areas.
Russell Schwartz is Professor of Biological Sciences and Computational Biology at Carnegie Mellon University, with courtesy appointments in Computer Science and Machine Learning. He serves as Head of the Ray and Stephanie Lane Computational Biology Department and Co-Director of the Joint CMU-University of Pittsburgh PhD Program in Computational Biology. His research develops computational models for complex biological systems, focusing on cancer genomics and macromolecular assembly. Key areas include reconstructing tumor clonal evolution through integration of bulk and single-cell sequencing data, developing algorithms for stochastic simulation of self-assembly systems, and creating optimization methods for fitting models to experimental data. Recent publications demonstrate advancements in single-cell lineage tracing, circulating tumor DNA analysis, and generative AI for biological databases. His team's BrM-Phylo software enables phylogeny reconstruction from matched transcriptome data in breast cancer brain metastases. NSF CAREER Award Alfred P. Sloan Research Fellowship Schwartz leads multiple NIH-funded projects on tumor evolution modeling and mentors students in computational cancer research. His laboratory develops open-source tools including FISHtrees for tumor phylogenetics and TUSV for integrating structural variants in clonal reconstruction.
Feras Saad is an Assistant Professor in the Computer Science Department at Carnegie Mellon University, affiliated with the Principles of Programming and Artificial Intelligence groups. He received his Ph.D. in Computer Science from the Massachusetts Institute of Technology (MIT) in 2022, where his dissertations on probabilistic programming systems earned him the George M. Sprowls PhD Thesis Award and Charles & Jennifer Johnson MEng Thesis Award. His research focuses on developing scalable computing systems for probabilistic modeling and inference, integrating ideas from programming languages and probabilistic AI. Key research themes include probabilistic programming languages, automated probabilistic model discovery, statistical estimation and testing, random sampling algorithms, and applications in science and engineering. His lab explores new techniques to improve reasoning systems through automation, accuracy, and scale. Dr. Saad has published extensively in top venues including PLDI, POPL, ICML, and Nature Communications. His work spans foundational computational questions to practical software systems for probabilistic inference. Research trends show consistent focus on bridging theoretical computer science with practical applications in probabilistic modeling, with recent advances in random sampling algorithms and probabilistic programming systems. Awards and honors include: George M. Sprowls PhD Thesis Award in Artificial Intelligence and Decision Making (2023) Charles & Jennifer Johnson MEng Thesis Award in Computer Science (2017) Editor's Highlight for Nature Communications paper (2024) He currently advises graduate students Gaurav Arya and Thomas Draper in the Probabilistic Computing Systems Lab. His research is supported by software libraries including GenSQL, BayesNF, and SPPL that enable practical applications across scientific domains.
Ruben Martins is an Assistant Research Professor and Master’s Program Director at the School of Computer Science, Carnegie Mellon University. He holds a Ph.D. from the Technical University of Lisbon, followed by postdoctoral research at the University of Oxford and UT Austin. His work focuses on constraint programming, program synthesis, and formal verification, with applications in improving programmer productivity and automating data science tasks. Education: Ph.D. in Computer Science, Technical University of Lisbon (2013) Postdoctoral Researcher, University of Oxford (2014-2015) Postdoctoral Researcher, UT Austin (2015-2017) Ruben's research bridges constraint programming with program synthesis, aiming to automate tasks such as vulnerability detection, code repair, and SQL synthesis. He developed Open-WBO , a MaxSAT solver that won multiple gold medals and is used in real-world optimization scenarios like seating arrangements for events. His work integrates large language models (LLMs) with traditional methods to enhance software analysis and debugging. Awards: Distinguished Paper Award at PLDI 2018 Gold Medal in MaxSAT Competitions for Open-WBO Advising & Grants: Ruben advises Master’s students and directs courses such as 15639, 15604, and others. His research has been supported by grants focusing on program synthesis and cybersecurity. Labs/Teams: Leads development of Open-WBO and contributes to projects like Crabtree (Rust API testing) and Pryde (evasion attack analysis).