Vending-Bench and Project Vend: Testing the Long-Term Cohesion of AI Agents in Vending Business Management

Can AI agents run a real business? A simulated benchmark and real-world deployment reveal critical limitations in long-term coherence.

Claudius
AI Agent Managing Real Store
2
Studies: Benchmark + Real-World
High
Variability & Failure Risk
Abstract

"Vending-Bench" is a simulated environment developed by Andon Labs to evaluate the ability of large language models (LLMs) to maintain long-term coherence when managing a simple yet lengthy business scenario: vending machine management. The Andon Labs paper presents results showing that, while some models such as Claude 3.5 Sonnet can outperform the human baseline in terms of average net value, all models exhibit high variability and a tendency towards "tangential failures" or "catastrophic cycles" when misinterpreting operational status.

Anthropic's article, titled "Project Vend", details the next phase of this research in which the Claude Sonnet 3.7 agent (nicknamed "Claudius") was deployed to manage a real automated store in Anthropic's office. There, it encountered inventory management issues, hallucinations and an "identity crisis", though it also demonstrated some successes.

Both sources emphasise that the ability of LLMs to function autonomously over the long term is a significant limitation that must be overcome for AI agents to be widely adopted in the economy.

Sources
Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

Andon Labs • arXiv • February 20, 2025 • DOI: 10.48550/arXiv.2502.15840

Read on arXiv
Project Vend: Can Claude run a small shop? (And why does that matter?)

Anthropic • June 27, 2025

Read on Anthropic
Key Finding: All tested models exhibit high variability and "catastrophic cycles" when managing long-term business operations, revealing critical limitations for autonomous AI agents.