Only What I Asked: Benchmarking AI Instruction Following

Developer akj1608 has introduced a new benchmark on Kaggle designed to test how well coding models adhere to specific, limited instructions. The project addresses a common frustration where AI assistants perform unrequested edits, such as fixing unrelated bugs or refactoring code when asked for a simple change. The benchmark consists of 22 tasks where models must apply a single edit while leaving the rest of the file untouched. Results indicate that while frontier models like Claude, GPT-6, and Gemini perform well, smaller or reasoning-heavy models often struggle with 'overreaching' or formatting issues. The study highlights that explicit warnings to 'change nothing else' do not always guarantee compliance and can sometimes negatively impact performance. The author concludes that surgical instruction-following remains a challenge, suggesting that future evaluations should focus on diff-based metrics rather than full-file rewrites to better measure model precision.
This is a summary. Read the full article at the original source:
Dev.toRelated stories
Promoting an online store in AI responses: why neural networks rarely link to product pages
The author conducted an experiment by asking Alice AI and Perplexity 14 typical consumer questions about electronics to analyze the link structure in…
Epilogue: There were many possibilities for things to go wrong
The author shares his experience of presenting a virtual AI-powered speaker at a major event in Blumenau, Brazil. Despite months of meticulous prepara…
Ollaya has emerged as a new platform designed to bring the ease of use associated with Ollama to the domain of open-source decision models. Inspired b…



