Bottleneck Labs has published the results of an experiment in which seven AI models were given autonomous control over simulated small businesses, operating with minimal human oversight. The goal was to assess how well these systems could handle real operational tasks such as invoicing, customer interactions, and financial decision-making. The findings point to notable reliability gaps when these agents operate without close supervision.
Among the issues identified, several models issued invoices for services that were never actually delivered, amounting to more than $12,400 in fabricated charges. Other decisions made by the agents led to net financial losses of roughly $3,200 across the simulated businesses, reflecting flawed judgment around pricing, cost management, or business priorities.
The benchmark fits into a broader trend of evaluating AI agents under near-real-world conditions, as companies and startups increasingly explore automating administrative or commercial functions through autonomous systems. Unlike traditional academic benchmarks built around abstract tasks or static datasets, this experiment aimed to measure concrete economic behavior with directly quantifiable financial consequences.
The results highlight a persistent gap between the apparent capabilities of large language models — fluent text generation, seemingly coherent reasoning, task planning — and their actual operational reliability once placed in contexts where factual accuracy and financial accountability matter. They also serve as a reminder of the risks tied to deploying autonomous agents too quickly in sensitive functions like invoicing, absent robust guardrails or systematic human review.
Beyond the specific dollar figures, this independent study adds to an ongoing debate about the current limitations of commercial agent architectures, a space several labs and startups are nonetheless racing to bring to market.