How I built the OenoBench wine benchmark with 3,266 questions using LLMs

The author shares their experience building OenoBench, a large-scale benchmark designed to evaluate how well language models understand wine. The project comprises 38,104 facts and 3,266 multiple-choice questions. Most of the work was automated: Claude Code wrote the code, five different LLMs generated the questions, and ten audit agents verified them, with API costs totaling around $800. The article details technical challenges, including issues with scrapers relying on model memory rather than external data, the methodology of question generation, and the analysis of errors made by LLM judges. The author also examines the performance of 16 models on the resulting leaderboard, concluding that the most complex questions in generated benchmarks are often the most flawed. This case study highlights both the potential and the limitations of using AI agents to create complex evaluation datasets.
This is a summary. Read the full article at the original source:
HabrRelated stories
Trump’s ‘Morally Binding’ AI ‘Accord,’ the Rise of AI Agents, and Extremists on the Ballot
This week’s episode of the Uncanny Valley podcast explores the evolving landscape of artificial intelligence and its intersection with politics. The d…
Google’s new Guided Vision feature can help you read the fine print
Google has officially launched its new Guided Vision feature within Gemini Live on compatible Android devices. This accessibility-focused tool leverag…
OpenAI has introduced new shopping-focused features for ChatGPT, enabling users to virtually try on clothing and accessories. By uploading their own p…



