Events
Project Match
FDS Data Science Project Match
Tuesday, September 1, 2026
2:00PM - 3:00PM
Location: Yale Institute for Foundations of Data Science, Kline Tower 13th Floor, Room 1327, New Haven, CT 06511 and via Webcast: https://yale.hosted.panopto.com/Panopto/Pages/Viewer.aspx?id=c06c3a2a-bc99-4903-96c3-b4b7014197c4
The FDS Data Science Project Match, hosted by the Yale Institute for Foundations of Data Science (FDS), is an opportunity for Yale faculty from any department or school within the university to connect with talented students from the departments of Statistics and Data Science, Applied Mathematics, and Computer Science. In a series of lightning-round talks, faculty will have exactly five minutes to pitch a current research problem, aiming to team up with students interested in tackling complex data challenges. This event facilitates collaboration on current research projects, offering a platform for faculty to present their data-driven initiatives and find skilled undergraduate and/or graduate students eager to contribute. It’s also a wonderful way to learn about the research of many Yale faculty.
MC: Brian Macdonald, Senior Lecturer and Research Scientist, Department of Statistics and Data Science
https://statistics.yale.edu/profile/brian-macdonald
|
Speaker: Joel Rozowsky (presenting on behalf of Mark Gerstein) (YSM) Research Scientist in Molecular Biophysics and Biochemistry joel.rozowsky@yale.edu Yale School of Medicine Talk Title: Genomics & Bioinformatics Research Opportunities in the Gerstein Lab The Gerstein lab conducts computational biology & bioinformatics research in the biomedical and genomic fields. We use various computational analytics methods including artificial-intelligence / machine-learning techniques to analyze large biomedical datasets and develop bioinformatics tools. The lab has particular focuses on the following areas of research: neurogenomics, personal genomes, genomic privacy and genome annotation. |
|
Speaker: Purushottam Dixit (SEAS) Assistant Professor of Biomedical Engineering purushottam.dixit@yale.edu Yale School of Engineering and Applied Science Talk Title: Reverse-Engineering the Optimal Traits of Microbiomes Using Minimax Entropy Microbial communities in the gut, soil, and ocean hold thousands of species, yet we still lack a simple theory for which species thrive where and what the community does. In the macroecology of plants and animals, measurable traits like leaf thickness or beak shape are found to predict species abundances. But in microbiomes, the physiological traits that govern growth cannot be measured organism by organism inside a living community. We instead plan to infer them directly from data. Using the minimax-entropy principle, a community’s composition is the least-biased distribution consistent with a few community-averaged traits, and the optimal traits are the ones that minimize the residual entropy, we reverse-engineer the traits of a microbiome directly from data. This compresses a community of thousands of taxa into a handful of interpretable low dimensional axes that predict composition, read out the environment, and forecast function, and we test the inferred traits against real genomes and metabolism to confirm they are genuine biology. https://sites.google.com/view/dixitlab/home |
|
Speaker: Quanquan Liu (Yale) Assistant Professor of Computer Science quanquan.liu@yale.edu Yale University Talk Title: Private Intelligence: Evaluating and Designing Local AI Models with the Tenstorrent TT-QuietBox 2 This project will evaluate the Tenstorrent TT-QuietBox 2 as a platform for running and adapting local AI models for privacy-sensitive tasks. We will benchmark open-weight language and multimodal models on document analysis, information extraction, question answering, summarization, and code generation, measuring accuracy, latency, throughput, memory use, energy consumption, and multi-device scalability. We will also develop smaller, task-specialized models using approaches that may interact with larger frontier models. The project will produce reproducible benchmarks, deployment tools, and practical guidance for determining when local AI can complement or replace cloud-based systems while reducing external exposure of sensitive data. |
|
Speaker: Soheil Ghili (SOM) Associate Professor of Marketing soheil.ghili@yale.edu Yale School of Management Talk Title: The Agentic Economy: When Do AI Agents Start Writing Contracts With Each Other? |
|
Speaker: Kexin Zhang (Yale) Assistant Professor of Molecular Biophysics and Biochemistry (appointment begins 9/1/2026) kexin.k.zhang@yale.edu Yale University Talk Title: Context-Aware and Adaptive Cryo-EM |
|
Speaker: Ibrahim Ibne Alam, MD (presenting on behalf of Leandros Tassiulas) (SEAS) Postdoctoral Associate, SmartNets Lab, Department of Electrical and Computer Engineering mdibrahimibne.alam@yale.edu Leandros Tassiulas, John C. Malone Professor of Electrical and Computer Engineering leandros.tassiulas@yale.edu Yale School of Engineering and Applied Science Talk Title: Statistical Modeling for Data-Scarce Power Systems: Market Price Forecasting and EV Demand Analysis in Electricity Markets This project tackles two related forecasting problems in the electricity sector. Both share a common core challenge: making reliable predictions from limited, sparse, or anomalous data. The first thread involves day-ahead market clearing price forecasting for European electricity markets (e.g. Greek Market). These markets often have a short historical record (~6 years). Part of that record is also distorted by structural anomalies tied to major regional political or energy crises. We have tested existing ML and deep learning models, including LASSO, PCA, RF, LGBM, MLP, and transformer-based models like TOTO and Chronos. These models perform on par with published benchmarks. However, they noticeably underperform relative to results reported for larger, more stable markets (e.g. NordPool). This gap is likely a direct consequence of the short and anomalous training window. The second thread involves EV charging demand prediction, using proprietary station-level data from an industry partner. Aggregate hourly demand can be forecasted accurately. However, station-level forecasts currently perform little better than random guessing. This is because many stations see intermittent usage, concentrated in just a few scattered hours across sparse active days. We’re looking for students with a strong statistics/math background to bring rigorous statistical thinking to both problems. For price forecasting, this means characterizing the structural break in the data. It also means building interpretability and uncertainty analysis, to help explain the performance gap versus peer markets. For EV demand, this means applying statistical methods suited to sparse, intermittent series. It also means developing a smarter station clustering approach. Our naive geographic clustering did not work, so the new approach should incorporate road type, urban/rural context, and altitude. Both threads offer a strong opportunity to apply statistical modeling to real, messy, industry data, with direct impact on power-sector decision-making. |
|
Speaker: Theo McKenzie (Yale ) Assistant Professor of Statistics and Data Science theo.mckenzie@yale.edu Yale University Talk Title: Geometric Structure of Gaussian Matrices Spectral methods use the eigenvalues and eigenvectors of a matrix to uncover structure in data. They form the basis of many widely used tools in data science. A Gaussian random matrix is one of the simplest models of pure noise: its entries are chosen randomly from a normal distribution, with no underlying signal or prescribed structure. Even in this setting, the eigenvectors can exhibit rich and highly predictable geometric patterns. In this project, we will investigate this geometry and identify features that arise as universal consequences of randomness itself. Understanding these patterns gives us a baseline for distinguishing genuine structure in data from structure that can arise purely from noise.
|
|
Speaker: Armita Nourmohammad (Yale) Associate Professor of Immunobiology and Biomedical Engineering Associate Director of Theoretical Sciences, Yale Center for Systems and Engineering Immunology (CSEI) armita.nourmohammad@yale.edu Yale University Talk Title: Biology Informed Machine Learning for Cellular Reprogramming Cells sense molecular cues from their environment and integrate them into decisions to proliferate, differentiate, or die. This project will develop interpretable generative machine-learning models, grounded in biological domain knowledge, to predict how molecular recognition achieves its specificity and how extracellular signals shape cell fate — using the adaptive immune system as a uniquely rich testbed. The resulting models will enable rational reprogramming of cell populations, from engineering therapeutic immune cells to improving tissue regeneration.
|
|
Speaker: Snigdha Jain (Yale) Assistant Professor, Section of Pulmonary, Critical Care, and Sleep Medicine snigdha.jain@yale.edu Yale University Talk Title: Improving Diagnostic Decision-Making for Weaning Sedation and Ventilator Support in the Intensive Care Unit Every day, patients in the ICU are kept sedated and on the mechanical ventilator longer than necessary – not because it’s medically required, but because clinicians lack a fast, reliable way to synthesize the dozens of data points (vital signs, medications, ventilator settings, labs) needed to determine readiness for “wake-up” trials. The result: unnecessary sedation, prolonged ventilation, and higher rates of delirium and long-term functional and cognitive decline leading to loss of independence. This project needs a computer science collaborator to build and validate an algorithm that processes structured, multi-domain EHR data – hourly, in near real-time – to flag when a patient meets clinical criteria for sedation and ventilator weaning trials. The core challenge is a time-series classification problem: integrating heterogeneous streams (hemodynamics, respiratory parameters, neuro status, medication records) across, then validating algorithm output against clinician-documented gold-standard labels. This is a rare chance to design an algorithm with a clear path to clinical deployment – feeding directly into a decision-support tool that will be piloted across multiple health systems once validated. https://medicine.yale.edu/profile/snigdha-jain/ |
|
Speaker: Kazuki Irie (SEAS) Assistant Professor Yale School of Engineering and Applied Science Talk Title: Metalearning Continual Compositional Reinforcement Learning Algorithms One of the most natural learning abilities of humans and other animals is the ability to accumulate skills from past experience and compose them to accomplish more complex tasks. While reinforcement learning in artificial intelligence has made progress in decomposing complex behavior into reusable skills in a top-down fashion, it remains challenging to learn such skills, bottom-up, incrementally from a stream of experience, retain them, and flexibly compose them to solve increasingly complex tasks. Given the difficulty of hand-designing algorithms for this kind of continual compositional learning, this project aims to design a framework to metalearn them in a data-driven fashion. |
|
Speaker: Fengzhuo Zhang (presenting on behalf of Zhuoran Yang) (Yale) Postdoc with Zhuoran Yang Assistant Professor of Statistics and Data Science zhuoran.yang@yale.edu Yale University Talk Title: Influence-Guided Self-Evolution: Scaling LLM Reasoning to Multi-Turn Agentic Tasks High-quality training data for advanced reasoning is scarce and expensive to produce. To address this bottleneck, a major open question in the field is whether Large Language Models (LLMs) can generate their own verifiable training data and independently train on it to “self-evolve.”
In our prior work, we explored this through an iterative co-training framework designed for math and coding QA tasks. In this setup, a Generator model drafts questions and reference answers from unstructured documents, while a Solver model trains on them. Instead of rewarding the Generator for simply creating “hard” questions, it is rewarded by an optimizer-aware influence score—a measure of how much a proposed question actually improves the Solver. Using Qwen3-8B models, this approach turned unstructured document pools into adaptive curricula, outperforming strong baselines by over 20% on Olympiad-level benchmarks.
For this project, our goal is to push the boundaries of self-evolution by expanding from single-turn QA to complex, multi-turn agentic tasks. We want to explore how influence-guided frameworks can help models self-improve in environments that require sequential decision-making, tool use, and long-horizon reasoning. Requirements: We are looking for students who are familiar with “vibe coding” or possess strong general coding expertise. Some empirical experience with deep learning, especially Pytorch with gnu is expected. A basic understanding of calculus, probability, and deep learning, along with a solid familiarity with linear algebra, is expected. Website: https://zhuoranyang.github.io
|
Add To: Google Calendar | Outlook | iCal File
- Project Match
Submit an Event
Interested in creating your own event, or have an event to share? Please fill the form if you’d like to send us an event you’d like to have added to the calendar.
