Drafts That Never Shipped: Do LLMs Treat Dead Proposals as Real Standards?

A new benchmark study explores whether Large Language Models (LLMs) mistakenly treat rejected or abandoned technical proposals—such as withdrawn emoji, deferred Python PEPs, and expired IETF drafts—as established standards. The research tested eight models from five different labs, including frontier models like Claude Sonnet 5 and Gemini 3.7 Flash. Findings reveal that models frequently hallucinate these 'dead' proposals as active, with 26% of answers asserting that non-shipped standards are real. Even when models correctly identify a proposal's status, they often fail to exclude it from generated code or citations, leading to non-compiling syntax or fake RFC references. The study highlights that while providing ground-truth evidence in prompts significantly reduces these errors, the issue remains a persistent challenge across both small and large models, suggesting that training data containing outdated specs is a primary source of these persistent technical hallucinations.
This is a summary. Read the full article at the original source:
Dev.toRelated stories
AI is creating a 'human premium' for art created by people
As generative artificial intelligence tools become increasingly capable of producing high-quality images, text, and music, a new economic trend is eme…
A recent exploration of Google’s AI-powered game development tools reveals the current limitations of generative technology in creative software desig…
AI Series Pilot for 900 Rubles on GPU with Kandinsky 6: Testing Character Dialogue
The author shares their experience creating a pilot episode for a personal series using the Kandinsky 6 model. The 52.2-second project features 19 sho…

