Technologies
Back
Artificial Intelligence & Machine Learning

I tested 36 AI models for fake packages and found zero

Dev.to
Advertisement468 × 90
I tested 36 AI models for fake packages and found zero

Developer Aarish Mansur has released a benchmarking tool designed to evaluate how frequently AI models hallucinate non-existent software packages. The benchmark, titled 'hallucinated-packages', analyzes code generated by LLMs by cross-referencing imported libraries against live PyPI and npm registries. In an initial test of 36 models, including various versions of Claude, Gemini, and GPT, the author observed a 0% hallucination rate for package names, suggesting that frontier models are generally reliable at identifying existing libraries. However, a follow-up pilot (Version 7) introduced more rigorous testing, including checks for non-existent functions within real packages. This iteration revealed that while models often avoid fake packages, they frequently hallucinate specific function calls within legitimate libraries. The author emphasizes that simple registry lookups are insufficient for comprehensive benchmarking and encourages the community to use these tools to further stress-test model reliability in real-world development scenarios.

This is a summary. Read the full article at the original source:

Dev.to
Advertisement468 × 90
Share
Artificial Intelligence & Machine Learning

Related stories

Advertisement970 × 250