Does compacting tool output lower a coding agent's API bill?

A recent experiment by the developers of the Torana framework investigated whether compacting tool output could reduce API costs for coding agents. Coding agents typically resend the entire conversation history, including verbose tool outputs, with every turn. The study tested two compaction methods—deterministic keyword filtering and model-gated summarization—against a control group using DeepSeek V4 Pro. Results showed that compaction provided negligible savings, with a median reduction of only about 2%. The primary reason for this outcome is the efficiency of modern prompt caching, which makes repeated context nearly free. Rewriting the prompt to include compacted data often invalidates the cache, leading to higher costs that negate the savings from reduced token counts. The author concludes that while compaction is ineffective under current caching economics, it may still hold value for providers with expensive cache reads or in environments where context limits are a bottleneck.
This is a summary. Read the full article at the original source:
Dev.toRelated stories
In a recent article on Habr, Selectel security engineer Anton explores the vulnerability of Large Language Models (LLMs) to various attacks, including…
What to do if you lose access to Claude while your workflow depends on it?
Since early October, Russian users have faced a new wave of Claude account blocks. For many companies, this has become a critical issue as the model w…
OpenAI has introduced 'Intelligent UI,' a new feature for ChatGPT that enables the AI to generate interactive visual interfaces directly within conver…



