首次评估印度多语言临床语音识别偏见,揭示模型在不同人群中的表现差异。
ASR Under the Stethoscope: Evaluating Biases in Clinical Speech Recognition across Indian Languages
- 对比多种主流模型在卡纳达语、印地语和印地英语上的表现
- 发现部分模型在方言或混码语境下准确率显著下降
- 暴露医生与患者、不同性别间系统性识别偏差,适合医疗AI公平性研究者
自动语音识别(ASR)正被用于记录临床对话,但在多语言且人口结构多元的印度医疗环境中其可靠性仍不明确。本研究首次对涵盖卡纳达语、印地语和印地英语的真实临床访谈数据进行系统性审计,对比了Indic Whisper、Whisper、Sarvam、Google Speech-to-Text、Gemma3n、Omnilingual、Vaani和Gemini等主流模型的性能。评估覆盖语言、说话人及人口子群体,重点关注患者与临床医生之间的错误模式,以及基于性别或交叉性因素的差异。结果揭示模型与语言间的显著性能差异,部分系统在印地英语上表现良好,但在混合语或本土语言中表现不佳。同时发现与说话人角色和性别相关的系统性差距,引发临床部署公平性的担忧。本研究构建了全面的多语言基准与公平性分析,强调发展文化与人口包容性更强的医疗语音识别系统的必要性。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diverse Indian healthcare contexts remains largely unknown. In this study, we conduct the first systematic audit of ASR performance on real world clinical interview data spanning Kannada, Hindi, and Indian English, comparing leading models including Indic Whisper, Whisper, Sarvam, Google speech to text, Gemma3n, Omnilingual, Vaani, and Gemini. We evaluate transcription accuracy across languages, speakers, and demographic subgroups, with a particular focus on error patterns affecting patients vs. clinicians and gender based or intersectional disparities. Our results reveal substantial variability across models and languages, with some systems performing competitively on Indian English but failing on code mixed or vernacular speech. We also uncover systematic performance gaps tied to speaker role and gender, raising concerns about equitable deployment in clinical settings. By providing a comprehensive multilingual benchmark and fairness analysis, our work highlights the need for culturally and demographically inclusive ASR development for healthcare ecosystem in India.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。