
In April 2026, we evaluated the latest and best-performing private (Gemini 3.1, OpenAI 5.4, Intron) and open source/sovereign deployable (Omnilingual, Gemma 4 26B, Kimi2.5) models for their performance on health-related audio questions in the languages commonly found in Pakistan and Nigeria. In our test, Gemini 3.1 Pro as the LLM + Intron or Meta's Omnilingual as the speech recognition model (ASR) scored most accurate, making the combination appropriate for WhatsApp and other async messaging based deployments, though likely too slow for usable voice-only services.
For Urdu, GPT-5.5 + Omni and Gemini 3.1 Flash-Lite + Omni both lead on accuracy at ~0.7, with Flash-Lite + Omni also the fastest among the top performers, and in general, at ~9s latency. Gemini 3.1 Pro Preview + Omni has fairly good accuracy but the highest latency (~32s), making it better suited to async use cases.
For Pashto, GPT-5.5 + Omni clearly leads on accuracy (~0.71) with moderate latency (~14s). Compared with the Urdu results, Pashto appears more accuracy-constrained overall: even the top model scores are lower than Urdu’s best performers, and the strongest real-time balance is likely Gemini 3.1 Flash-Lite + Omni, which keeps latency under ~8.5s while reaching mid-range accuracy.
For Sindhi, GPT-5.5 + Omni is the most accurate option (~0.73) at roughly 12s latency, with Gemini 3.1 Pro Preview and Gemini 3.1 Pro Preview + Omni close behind but slightly slower. Compared with Urdu, Sindhi’s top accuracy is lower, but it looks closer to Pashto in overall performance; Gemini 3.1 Flash-Lite + Omni again offers the practical speed–accuracy middle ground.

For Hausa, Gemini 3.1 Pro achieved the highest Semantic Similarity Score (Accuracy) while maintaining lower latency than most tested configurations (0.71; ~16.3s). Claude Sonnet 5 + Omni was fastest overall, but its accuracy was considerably lower (0.5; ~13.6s).

For Nigerian Fulfulde, Gemini 3.7 Flash + Omni achieved the highest Semantic Similarity Score (Accuracy) while delivering nearly the lowest latency, making it the strongest overall configuration (0.59; ~15.5s). Claude Sonnet 5 + Omni was marginally faster, but its accuracy was substantially lower (0.23; ~15.3s).

For Igbo, Kimi K3 + Intron achieved the highest Semantic Similarity Score (Accuracy) (0.55; ~23s), but GPT 5.6 Sol + Omni followed closely while delivering the lowest latency (0.51; ~14.91s). This made GPT 5.6 Sol + Omni the strongest overall speed–accuracy trade-off.
For Yoruba, Gemini 3.1 Pro achieved the highest Semantic Similarity Score (Accuracy) by a clear margin while maintaining moderate latency (0.61; ~17.2s). Claude Sonnet 5 + Omni was fastest overall, but its accuracy was substantially lower (0.27; ~13.9s).
For Nigerian Pidgin, Kimi K3 + Intron achieved the highest Semantic Similarity Score (Accuracy) (0.75; ~22.2s), while Kimi K3 + Omni followed closely with lower latency (0.71; ~18.8s). Claude Sonnet 5 + Omni was fastest overall, but its accuracy was considerably lower (0.62; ~16.2s).
In December, 2025 We tested new LLM+ASR AI workflows to understand real-world deployment Swahili, Kikuyu, and Kinyarwanda voice applications. Fine-tuned African ASR models + modern LLMs (GPT-5.1, Gemini 3) achieved 85-96% accuracy with 4-10 second response times—fast enough for phone-based services. In short, Voice AI for many African languages is now technically viable for WhatsApp and basic phone deployments at scale.
Read the PaperFor Kinyarwanda, DeepSeek V4 + Omni achieved the highest Semantic Similarity Score (Accuracy) while maintaining relatively low latency (0.92; ~7.1s). Kimi K3 + Omni was fastest overall, but its accuracy was noticeably lower (0.78; ~6.7s).
For Swahili, Gemini 3.1 Pro was the clear leader for Swahili, achieving both the highest semantic accuracy and lowest response latency (0.95; ~5.8s). It delivered the strongest speed–accuracy balance of every model configuration tested.
For Kikuyu, Gemini 3.1 Pro + Kikuyu Whisper v2 was the most accurate by a clear margin (0.86; ~9.5s). Claude Sonnet 5 + Kikuyu Whisper v2 was fastest overall (0.63; ~6.3s), while Gemini 3.7 Flash + Kikuyu Whisper v2 offered a strong balance of speed and accuracy (0.82; ~7.1s).



.png)





.png)
