TypeSafe Jev Played Chess — And Landed Next to Reasoning Models

Developer Maxim Saplin has integrated TypeSafe’s Jev model into his LLM Chess leaderboard to test its performance against traditional chat-based LLMs. Unlike standard models that generate free-form text, Jev is a 'System One' model designed for structured classification tasks, where it selects from a predefined set of options. By treating chess moves as a classification problem, Saplin successfully enabled Jev to play full games. The results were striking: Jev achieved an Elo rating comparable to mid-tier reasoning models while being significantly faster and cheaper, costing only $0.0015 per game. While Jev struggles with generative tasks like spelling, its efficiency in decision-making and protocol adherence makes it a compelling tool for structured agentic workflows. The experiment highlights that specialized, non-chat models can compete with larger reasoning models in specific, constrained environments, offering a high-performance alternative for developers building automated decision systems.
This is a summary. Read the full article at the original source:
Dev.toRelated stories
As enterprises increasingly integrate autonomous AI agents into their workflows, a significant financial risk has emerged: unbounded consumption. Acco…
Stopping AI’s Runaway Dangers Will Take More Than Just Talk About P(doom)
In a recent guest column for CNET, author Jamie Bartlett explores the escalating risks associated with advanced artificial intelligence. Bartlett argu…
OpenAI forms math advisory group as its AI resolves more than 100 open problems
OpenAI has officially established a dedicated mathematical advisory group to oversee its ongoing research into advanced AI reasoning. This development…



