Technologies
Back
Artificial Intelligence & Machine Learning

Only What I Asked: Benchmarking AI Instruction Following

Dev.to
Advertisement468 × 90
Only What I Asked: Benchmarking AI Instruction Following

Developer akj1608 has introduced a new benchmark on Kaggle designed to test how well coding models adhere to specific, limited instructions. The project addresses a common frustration where AI assistants perform unrequested edits, such as fixing unrelated bugs or refactoring code when asked for a simple change. The benchmark consists of 22 tasks where models must apply a single edit while leaving the rest of the file untouched. Results indicate that while frontier models like Claude, GPT-6, and Gemini perform well, smaller or reasoning-heavy models often struggle with 'overreaching' or formatting issues. The study highlights that explicit warnings to 'change nothing else' do not always guarantee compliance and can sometimes negatively impact performance. The author concludes that surgical instruction-following remains a challenge, suggesting that future evaluations should focus on diff-based metrics rather than full-file rewrites to better measure model precision.

This is a summary. Read the full article at the original source:

Dev.to
Advertisement468 × 90
Share
Artificial Intelligence & Machine Learning

Related stories

Advertisement970 × 250