Publications

Research on long-horizon LLM agents, post-training, data attribution, model safety, and AI for science.

2026 · arXiv preprint

Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

Yijun Pan, Yukun Lian, Kunyu Shi, Junbo Li, Hongwei Xue, Sicong Xie, Guannan Zhang, Xiaoying Xing

A long-horizon marketplace for measuring how LLM agents create or lose value through sourcing, pricing, inventory, service, compliance, and capital decisions.

2026 · arXiv preprint

FlexRec: Adapting LLM-based Recommenders for Flexible Needs via Reinforcement Learning

Yijun Pan, Weikang Qiu, Qiyao Ma, Mingxuan Ju, Tong Zhao, Neil Shah, Rex Ying

A reinforcement-learning framework with item-level counterfactual rewards and uncertainty-aware scaling for adaptable LLM recommenders.

2025 · NeurIPS Datasets and Benchmarks Track

DATE-LM: Benchmarking Data Attribution Evaluation for Large Language Models

Cathy Jiao, Yijun Pan, Emily Xiao, Daisy Sheng, Niket Jain, Hanzhang Zhao, Ishita Dasgupta, Jiaqi W. Ma, Chenyan Xiong

A unified benchmark and public leaderboard for evaluating LLM data-attribution methods across selection, safety, and factual attribution.

2025 · arXiv preprint

Detecting and Filtering Unsafe Training Data via Data Attribution with Denoised Representation

Yijun Pan, Taiwei Shi, Jieyu Zhao, Jiaqi W. Ma

A targeted attribution method for identifying harmful training samples and reducing unsafe model behavior after filtering and retraining.

2024 · arXiv preprint

Bridging AI and Science: Implications from a Large-Scale Literature Analysis of AI4Science

Yutong Xie*, Yijun Pan*, Hua Xu, Qiaozhu Mei (* Equal Contribution)

A large-scale map of the gaps between scientific problems and the AI methods currently used to address them.

2024 · NeurIPS Datasets and Benchmarks Track · Spotlight

dattri: A Library for Efficient Data Attribution

Junwei Deng, Ting-Wei Li, Shiyuan Zhang, Shixuan Liu, Yijun Pan, Hao Huang, Xinhe Wang, Pingbang Hu, Xingjian Zhang, Jiaqi W. Ma

An open-source library that unifies efficient data-attribution methods, utilities, and reproducible benchmarks behind a common API.

2024 · MSLD

Interpreting Spatial Reasoning Capabilities in Language Models

Yijun Pan*, Sushrita Rakshit*, Daniel Tian*, Hua Shen, Kenan Alkiek, David Jurgens (* Equal Contribution)

An interpretability study of where and how spatial reasoning capabilities emerge inside language models.