Technologies
Back
Artificial Intelligence & Machine Learning

I Surveyed 123 People in India to Benchmark Frontier AI

Dev.to
Advertisement468 × 90
I Surveyed 123 People in India to Benchmark Frontier AI

A new benchmarking study has evaluated how frontier AI models handle complex institutional dilemmas in India. By surveying 123 university students, the researcher created a dataset of 24 real-world disputes across education, healthcare, justice, and finance. The study tested models including Gemini, Claude, Qwen, and DeepSeek, focusing on metrics like ambiguity resistance, demographic parity, and sycophancy. Findings reveal a significant 'reasoning bias paradox,' where models with extended thinking capabilities often generated rationalizations that led to higher rates of demographic decision-flipping. While some models excelled at maintaining consistency in financial scenarios, others struggled with justice-related tasks. The research highlights that despite advancements, demographic parity remains an unsolved challenge. The author provides a comprehensive leaderboard on Kaggle, encouraging further investigation into regional language bias and the impact of legal precedent grounding on model performance.

This is a summary. Read the full article at the original source:

Dev.to
Advertisement468 × 90
Share
Artificial Intelligence & Machine Learning

Related stories

Advertisement970 × 250