Bengaluru | 17 September 2026: Humyn Labs, a Physical AI research lab focused on accelerating robot deployment‑readiness, has launched the second edition of BRIDGE, its global benchmark report that exposes the gap between human speech and AI voice models.
The study evaluated 23 voice AI models across 23 languages in real‑world noisy conversations, mapping performance across accents, dialects, overlapping speech, and conversational density. Models tested include Sarvam v3, Gemini 3 Pro, ElevenLabs, among others.
Key Findings
- Overlapping speech raised average error rates from 41.2% to 45.2%.
- Dialect gaps proved significant: Bengali standard scored 42.4%, while a regional dialect outside Kolkata scored 51.0%. Spanish showed similar divergence, with Argentinian Spanish at 7.85% vs Venezuelan Spanish at 16.04%.
- Model choice matters: ElevenLabs averaged 5.8% error, compared to 24.6% for GPT‑4o‑mini‑transcribe, a 4x gap on identical audio.
- Silence problem replicated globally: Brazilian Portuguese calls showed 18.8% error for pauses over 150 seconds vs 12.4% for shorter gaps.
- Error types differ structurally:
- Substitution errors (wrong word) dominated 19/23 models.
- Omission errors (dropping words) were common in OpenAI, Speechmatics, and Gnani Vachana.
- Fabrication errors (invented words) were seen in Gemini Flash, adding 9.3% fabricated content.
Manish Agarwal, Co‑Founder, Humyn Labs: “Voice is a critical interface for Physical AI. If models cannot handle overlapping speech, interruptions, code‑switching, pauses, and linguistic diversity, the gap impacts customer experience, automation, and trust. BRIDGE helps builders evaluate whether models are truly ready to scale.”
Ishank Gupta, Co‑Founder, Humyn Labs: “Physical AI cannot learn the real world through vision alone. Sound carries information about people, actions, distance, and intent. BRIDGE provides the missing evaluation layer for speech models, testing them across conditions that mirror real‑world environments.”
Benchmark Methodology
- Seven core metrics: overlapping speech, conversational density, code‑switching, pauses, dialects, and more.
- Coverage includes Indic languages, Latin American Spanish, Brazilian Portuguese, and Vietnamese.
- Built on 200+ hours of human‑verified audio, collected across two to three districts per language.
- Distinguishes script mismatch vs genuine transcription error, revealing commercial implications for workflow design.
Commercial Impact
- Best single model achieved 10.7% loanword‑adjusted error.
- Best per language reached 9.7%, while theoretical best per call achieved 8.9%.
- One model won 78.6% of files outright, underscoring the importance of model routing economics.

About Humyn Labs
Humyn Labs is a Physical AI research engine pioneering source‑first data collection, labelling, and enrichment across voice, vision, motion, and touch. Its mission is to accelerate robot deployment‑readiness by building multi‑modality enriched signals that help machines interpret the complexity of the real world. Operating across India, Southeast Asia, LATAM, and the Middle East, Humyn Labs combines human intelligence, technology, and deep research to shape the intelligence layer for inclusive, accurate, and scalable AI.
About BRIDGE
BRIDGE is a global independent ASR benchmark evaluating commercial and open‑source models on field‑collected, human‑verified conversational audio. Using a seven‑metric scoring stack, BRIDGE maps the gap between human speech and AI voice models, covering Indic languages, Latin American Spanish, Brazilian Portuguese, and Vietnamese, built on 200+ hours of real‑world audio.


Aviva Life Insurance Launches Term Insurance with Return of Premiums Proposition
PM Sports Welcomes Kabaddi Star Naveen Kumar as Brand Ambassador