Cents Matter: Does a Model Notice When One Financial Fact Changes?

A new benchmark, 'Cents Matter / Centavos Importam,' evaluates how well LLMs handle precise financial reasoning. Created for the Kaggle Benchmarking Challenge, the test consists of 40 synthetic cases in Brazilian Portuguese, organized into 20 minimal pairs where a single material fact is altered. The benchmark requires models to correctly identify discrepancies in monetary calculations without external tools. The study tested Gemini 3.7 Flash, gpt-oss-20b, and Claude Haiku 4.5. Results showed significant performance gaps: Gemini 3.7 Flash achieved 97.5% accuracy, while Claude Haiku 4.5 struggled with consistency, scoring 47.5%. The author emphasizes that while these models excel at general tasks, their ability to maintain logical consistency when financial variables shift is inconsistent. The benchmark is open-source, providing a transparent look at how different architectures handle deterministic financial logic, highlighting that high-level reasoning does not always guarantee precision in accounting-style arithmetic.
This is a summary. Read the full article at the original source:
Dev.toRelated stories
Tiny model. Big decisions. — How I built a 144M-parameter typed decision model that routes 82% of agent decisions off LLMs
Developer Perry Link has introduced Phocinae-Largha-150M-v1, a specialized 144-million parameter model designed to handle repetitive decision-making t…
“Software is over”: Bold AI developer takes aim at Adobe with open source clones
Developer Brandon Thomas has launched an ambitious project to replace Adobe’s Creative Suite with a suite of seven open-source applications. Utilizing…
Google launches Playground, an AI-powered browser game creation platform
Google has introduced Playground, an experimental platform designed to simplify browser game development through generative AI. The tool allows users…



