针对印度多语言临床语音识别的偏见问题,提出统一去偏方法SamaVaani。
SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages

- 系统审计8个模型在卡纳达语、印地语和印式英语上的表现
- 发现语音识别在地区性语言中准确率显著下降,且存在性别与角色偏差
- 提出SamaVaani去偏技术,兼顾性能提升与公平性改善
自动语音识别(ASR)在临床记录中应用日益广泛,但在印度多语言、人口多元的医疗环境中其可靠性仍不明确。本研究对涵盖卡纳达语、印地语和印式英语的真实精神科访谈数据,系统评估了八种前沿模型(IndicWhisper、WhisperLargeV3、Sarvam、GoogleS2T、Gemma3n、OmniLingual、Vaani、Gemini)的表现。结果表明各模型在不同语言间差异显著,部分系统在印式英语中表现良好,但在地方语言中表现不佳。进一步对表现最佳的两个开源模型(Gemma3n 和 OmniLingual)采用多种微调策略后,发现识别误差与说话人角色及性别存在系统性关联,提示临床部署存在公平性风险。通过公平性感知微调可有效缓解该问题。为此,本文提出SamaVaani——一种统一的去偏技术,能同时提升识别性能并增强跨人群公平性。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diverse Indian healthcare context remains largely unknown. In this study, we first conduct the systematic audit of ASR performance on real-world psychiatric interview data spanning Kannada, Hindi and Indian English, comparing eight state-of-the-art models including IndicWhisper, WhisperLargeV3, Sarvam, GoogleS2T, Gemma3n, OmniLingual, Vaani, and Gemini. Our results reveal substantial variability across models and languages, with some systems performing competitively in Indian English but failing in regional speech. We further fine-tune two of the best performing opensource models, i.e., Gemma3n and OmniLingual, using various methods. With this, we uncover systematic performance gaps tied to speaker role and gender, raising concerns about equitable deployment in clinical settings, which are further mitigated by fairness-aware fine-tuning. To this end, we propose SamaVaani, a unified debiasing technique that simultaneously improves ASR performance and improves fairness across demographic groups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。