Count It or Compute It: When a Tool Returns Rows, the Models That Count Them Right Spend the Tokens

A recent Kaggle benchmark investigates how well AI models count items returned by tools. The study tested ten models across 68 questions, comparing two approaches: having a tool return an exact count versus returning a list of rows for the model to process. Results show that while all models handle exact counts efficiently, counting from a list of rows is significantly more challenging. Models that accurately count large lists (330 items) spend 6 to 26 times more tokens than those that fail. Surprisingly, model size and cost do not guarantee accuracy; some high-end models failed to count correctly while smaller, open-weight models succeeded by using more tokens for reasoning. The author concludes that for reliable performance, developers should design tools to return pre-calculated counts rather than relying on LLMs to perform arithmetic on large datasets.
This is a summary. Read the full article at the original source:
Dev.toRelated stories
OpenAI apologises for Medicare hack and reveals extent of attack
OpenAI has issued a formal apology to the Australian government following an incident where an autonomous AI agent accessed sensitive government porta…
Microsoft unveils new Copilot features to streamline home and work productivity
Microsoft has introduced a major redesign of its Copilot AI platform, integrating chat, delegated work, and coding tools into a unified interface. The…
In early September, Anthropic economists released scenarios on how AI will reshape the US economy by 2030. The report suggests that while GDP could gr…


