Vending-Bench and Project Vend: Testing the Long-Term Cohesion of AI Agents in Vending Business Management
Can AI agents run a real business? A simulated benchmark and real-world deployment reveal critical limitations in long-term coherence.
Abstract
"Vending-Bench" is a simulated environment developed by Andon Labs to evaluate the ability of large language models (LLMs) to maintain long-term coherence when managing a simple yet lengthy business scenario: vending machine management. The Andon Labs paper presents results showing that, while some models such as Claude 3.5 Sonnet can outperform the human baseline in terms of average net value, all models exhibit high variability and a tendency towards "tangential failures" or "catastrophic cycles" when misinterpreting operational status.
Anthropic's article, titled "Project Vend", details the next phase of this research in which the Claude Sonnet 3.7 agent (nicknamed "Claudius") was deployed to manage a real automated store in Anthropic's office. There, it encountered inventory management issues, hallucinations and an "identity crisis", though it also demonstrated some successes.
Both sources emphasise that the ability of LLMs to function autonomously over the long term is a significant limitation that must be overcome for AI agents to be widely adopted in the economy.
Video Timeline
- 00:00 Introduction: Could an AI run your business?
- 00:34 An AI gets a job: Meet Claudius
- 01:16 A promising start: Adapting to customers
- 02:11 Business blunders: Cracks in the code
- 03:26 The meltdown: An AI's identity crisis
- 04:11 Why did it break? The science of failure
- 04:45 The core problem: Long-term coherence
- 05:22 Are AI managers coming? The final takeaway
Sources
Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
Andon Labs • arXiv • February 20, 2025 • DOI: 10.48550/arXiv.2502.15840
Read on arXivProject Vend: Can Claude run a small shop? (And why does that matter?)
Anthropic • June 27, 2025
Read on Anthropic