Tiny model. Big decisions. — How I built a 144M-parameter typed decision model that routes 82% of agent decisions off LLMs

Developer Perry Link has introduced Phocinae-Largha-150M-v1, a specialized 144-million parameter model designed to handle repetitive decision-making tasks for AI agents. Unlike generative models, this encoder-based model focuses on typed outputs—such as yes/no or selection tasks—providing deterministic results with low latency (18.6ms on GPU). By acting as a 'System 1' gate, the model can process routine agent queries locally, escalating only ambiguous cases to larger, more expensive LLMs. This approach successfully reduced LLM API calls by 82% in testing while maintaining high accuracy. The project is open-source under the Apache-2.0 license, emphasizing transparency in metrics like calibration and option-order stability. It is intended for developers looking to optimize agentic workflows by offloading simple, repetitive logic from heavy generative models to a lightweight, efficient, and cost-effective alternative.
This is a summary. Read the full article at the original source:
Dev.toRelated stories
Cents Matter: Does a Model Notice When One Financial Fact Changes?
A new benchmark, 'Cents Matter / Centavos Importam,' evaluates how well LLMs handle precise financial reasoning. Created for the Kaggle Benchmarking C…
“Software is over”: Bold AI developer takes aim at Adobe with open source clones
Developer Brandon Thomas has launched an ambitious project to replace Adobe’s Creative Suite with a suite of seven open-source applications. Utilizing…
Google launches Playground, an AI-powered browser game creation platform
Google has introduced Playground, an experimental platform designed to simplify browser game development through generative AI. The tool allows users…



