I tested 36 AI models for fake packages and found zero

Developer Aarish Mansur has released a benchmarking tool designed to evaluate how frequently AI models hallucinate non-existent software packages. The benchmark, titled 'hallucinated-packages', analyzes code generated by LLMs by cross-referencing imported libraries against live PyPI and npm registries. In an initial test of 36 models, including various versions of Claude, Gemini, and GPT, the author observed a 0% hallucination rate for package names, suggesting that frontier models are generally reliable at identifying existing libraries. However, a follow-up pilot (Version 7) introduced more rigorous testing, including checks for non-existent functions within real packages. This iteration revealed that while models often avoid fake packages, they frequently hallucinate specific function calls within legitimate libraries. The author emphasizes that simple registry lookups are insufficient for comprehensive benchmarking and encourages the community to use these tools to further stress-test model reliability in real-world development scenarios.
This is a summary. Read the full article at the original source:
Dev.toRelated stories
Can ‘super intelligence’ and a non-binding safety pact solve AI’s image problem?
In a recent episode of the Equity podcast, the discussion centered on the Trump administration's strategic efforts to rebrand artificial intelligence.…
Developer Builds AI 'Warden' to Resolve Hostel Roommate Disputes
To address recurring conflicts over badminton court bookings in his college hostel, developer Achintya Singh created 'Hostel Nexus,' an AI-powered dis…
Developer Builds AI-Powered Scam Detector for Family Using Open-Weight Gemma
A developer has created 'Chithi,' an open-source tool designed to protect non-English speaking family members from financial scams. The application us…



