Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
A long-horizon marketplace for measuring how LLM agents create or lose value through sourcing, pricing, inventory, service, compliance, and capital decisions.
Research on long-horizon LLM agents, post-training, data attribution, model safety, and AI for science.
A long-horizon marketplace for measuring how LLM agents create or lose value through sourcing, pricing, inventory, service, compliance, and capital decisions.
A reinforcement-learning framework with item-level counterfactual rewards and uncertainty-aware scaling for adaptable LLM recommenders.
A unified benchmark and public leaderboard for evaluating LLM data-attribution methods across selection, safety, and factual attribution.
A targeted attribution method for identifying harmful training samples and reducing unsafe model behavior after filtering and retraining.
A large-scale map of the gaps between scientific problems and the AI methods currently used to address them.
An open-source library that unifies efficient data-attribution methods, utilities, and reproducible benchmarks behind a common API.
An interpretability study of where and how spatial reasoning capabilities emerge inside language models.