12 of 13 AI models knew the new name and still wrote the old one

A recent benchmarking study on Kaggle, titled 'Semconv Drift,' tested 17 AI models on their ability to generate accurate OpenTelemetry instrumentation code. The results reveal a significant discrepancy between model knowledge and output: 12 out of 13 models correctly identified updated attribute names when asked directly, yet still utilized retired, deprecated names when generating code. The study highlights that models often prioritize common, outdated patterns found in training data over current specifications. Even when prompted with version-specific instructions, models showed only a 36% reduction in errors, sometimes inventing non-existent names instead. The author concludes that relying on AI for technical tasks requires rigorous validation, as models can produce 'half-fixes'—such as updating units while keeping deprecated names—that pass code reviews but break telemetry pipelines. The findings suggest that developers should implement automated linting against official specifications rather than trusting AI-generated code blindly.
This is a summary. Read the full article at the original source:
Dev.toRelated stories
The author shares their personal experience in optimizing the editing process for short vertical videos for Reels and Shorts. Previously, processing h…
Local Zoom Assistant: Experience with NLI Model Integration
The author continues a series of articles on developing a local assistant for video conferencing. The tenth installment examines the practical experie…
Anthropic restricts internet access for internal AI evaluations
Anthropic has announced a significant change to its safety protocols by cutting off internet access for all internal AI model evaluations. This decisi…


