Technologies
Back
Artificial Intelligence & Machine Learning

Whisper Keeps Correcting Nigerian Speech. Here's How I Measured It

Dev.to
Advertisement468 × 90
Whisper Keeps Correcting Nigerian Speech. Here's How I Measured It

Developer Nadine V. has released a benchmark study investigating how OpenAI's Whisper speech-to-text models handle Nigerian English and Pidgin. The research highlights a significant issue: standard models often 'correct' dialect-specific vocabulary into standard English, leading to inaccurate transcriptions. By comparing zero-shot performance of Whisper large-v3 and small models against a custom fine-tuned version, the study demonstrates that fine-tuning can reduce the Word Error Rate (WER) by approximately 80%. The author developed an auto-classifier using 120 dialect-specific rules to distinguish between genuine recognition errors and intentional dialect normalization. The findings suggest that standard metrics often penalize models for correctly transcribing local dialects. The project, which includes a detailed methodology and open-source scoring scripts, emphasizes the importance of evaluating how models handle linguistic diversity rather than relying solely on traditional error rates. The full analysis and datasets are available on GitHub and Kaggle.

This is a summary. Read the full article at the original source:

Dev.to
Advertisement468 × 90
Share
Artificial Intelligence & Machine Learning

Related stories

Advertisement970 × 250