About Me
Building LLM agents for the messiness of the real world.
I study long-horizon decision-making under sparse, noisy, and delayed feedback. I build simulation environments and develop post-training methods to make agents more capable, reliable, and adaptive in real-world settings.
I am pursuing a two-year M.S. in Computer Science at Yale University. I hold dual bachelor’s degrees in Data Science from the University of Michigan and Electrical and Computer Engineering from Shanghai Jiao Tong University.
Experience
Alibaba Inc.
Research Intern · AliStar Top Talent Program
Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
We build a long-horizon, realistic business world to measure whether agents can operate end-to-end businesses.
Snap Inc. & Yale University
Student Researcher
FlexRec: Adapting LLM-based Recommenders for Flexible Needs via Reinforcement Learning
Real user needs are diverse, and one item can carry different value when the need changes. We introduce FlexRec, a post-training framework that aligns LLM recommenders across multiple needs.
Carnegie Mellon University
Machine Learning Research Intern
DATE-LM: Benchmarking Data Attribution Evaluation for Large Language Models
Data attribution promises to explain how training data shapes LLM behavior, but existing evaluations are fragmented. We introduce DATE-LM, a unified benchmark for comparing attribution methods across practical LLM applications.
UIUC & USC
Machine Learning Research Intern
Detecting and Filtering Unsafe Training Data via Data Attribution with Denoised Representation
Small amounts of unsafe training data can meaningfully change model behavior, while fixed moderation categories can miss emerging risks. We use data attribution to connect harmful behavior back to influential training examples and support targeted filtering.
University of Michigan & UIUC
Student Researcher
dattri: A Library for Efficient Data Attribution
Data attribution methods are useful but difficult to implement and compare consistently. We introduce dattri, an open-source PyTorch library for developing, benchmarking, and deploying efficient data-attribution methods.
University of Michigan
Student Researcher
Bridging AI and Science: Implications from a Large-Scale Literature Analysis of AI4Science
AI and scientific research are advancing quickly, but useful methods and real scientific needs do not always meet. We map the AI4Science literature at scale to reveal these gaps and surface opportunities for cross-disciplinary collaboration.