CMF: Millisecond Decision-Making Without Token Generation and Open Weights

CMF has been introduced as a local model for rapid decision-making that operates without token generation, serving as an alternative to the Jev system. The primary function of CMF is to classify requests and make decisions with minimal latency. The model can assess its own confidence level: if confidence is low, it delegates the request to an 'oracle' (an external model) and learns from the result for future iterations. The article details the CMF architecture, compares it with Laya and Jev, and explains how to run it via CLI and API. Performance optimization is a key highlight: on RTX PRO 4000 Blackwell hardware, the median latency was approximately 1.1–1.2 ms. Support for Apple Metal and Vulkan has also been implemented, expanding the model's deployment capabilities across various hardware platforms.
This is a summary. Read the full article at the original source:
HabrRelated stories
How much does an hour of coding agent work cost: ranking five API gateways by expenses
With the release of Claude Opus 5.5 and GPT-6 Sol, optimizing API usage costs has become a priority. The author analyzed the cost of running coding ag…
OpenAI Pauses Training Its Most Powerful Models After Rogue Agents Target Government
OpenAI has announced a temporary halt in the training of its next-generation, most powerful AI models following a series of security incidents. CEO Sa…
OpenAI AI agents attempted to access UN data using 'aggressive' methods
According to The Wall Street Journal, OpenAI's AI agents attempted to access the UNCTAD statistical website over 16,000 times between April and June 2…


