I Surveyed 123 People in India to Benchmark Frontier AI

A new benchmarking study has evaluated how frontier AI models handle complex institutional dilemmas in India. By surveying 123 university students, the researcher created a dataset of 24 real-world disputes across education, healthcare, justice, and finance. The study tested models including Gemini, Claude, Qwen, and DeepSeek, focusing on metrics like ambiguity resistance, demographic parity, and sycophancy. Findings reveal a significant 'reasoning bias paradox,' where models with extended thinking capabilities often generated rationalizations that led to higher rates of demographic decision-flipping. While some models excelled at maintaining consistency in financial scenarios, others struggled with justice-related tasks. The research highlights that despite advancements, demographic parity remains an unsolved challenge. The author provides a comprehensive leaderboard on Kaggle, encouraging further investigation into regional language bias and the impact of legal precedent grounding on model performance.
This is a summary. Read the full article at the original source:
Dev.toRelated stories
A frontend developer has created a personalized AI-driven command center designed to streamline daily productivity and personal goal management. By le…
AI chatbots may be causing a global 'knowledge collapse' by reducing information diversity
A new study from the University of Copenhagen suggests that AI chatbots could be shrinking human knowledge by providing less diverse information compa…
Jev and Laya beyond the hype: what decision models do that LLMs don't
Developer Tiago Vilas Boas explores the practical limitations of using Large Language Models (LLMs) for every task in an agentic workflow. While LLMs…



