音频大模型在临床决策中会受声音特征影响,导致推荐差异高达35%。
MedVoiceBias: A Controlled Study of Audio LLM Behavior in Clinical Decision-Making
- 用36种语音合成170个病例,对比音频与文本输入的决策差异。
- 音频输入使手术建议最多减少35%,老年声音比年轻声音少12%推荐。
- 情绪识别差导致无法检测情绪偏见,需设计抗偏见模型。
随着大型语言模型从文本转向语音交互,其在临床场景中可能因语音中的副语言线索引入新漏洞。我们在170个临床病例上进行了评估,每个病例由36种不同年龄、性别和情绪的语音特征合成。结果显示严重模态偏差:音频输入的手术建议与相同文本输入相比最高相差35%,某一模型甚至减少了80%的建议。进一步分析发现,年轻与老年语音间的推荐差距可达12%,且多数模型即使采用思维链提示也未能消除该差异。尽管显式推理可消除性别偏见,但因情绪识别性能不佳,情绪影响未被检测到。这些结果表明,音频大模型会依据患者声音特征而非医疗证据做出临床决策,存在加剧医疗不平等的风险。我们得出结论:在临床部署前,亟需构建具备偏见感知能力的模型架构。
原文摘要 · Abstract (English)
As large language models transition from text-based interfaces to audio interactions in clinical settings, they might introduce new vulnerabilities through paralinguistic cues in audio. We evaluated these models on 170 clinical cases, each synthesized into speech from 36 distinct voice profiles spanning variations in age, gender, and emotion. Our findings reveal a severe modality bias: surgical recommendations for audio inputs varied by as much as 35% compared to identical text-based inputs, with one model providing 80% fewer recommendations. Further analysis uncovered age disparities of up to 12% between young and elderly voices, which persisted in most models despite chain-of-thought prompting. While explicit reasoning successfully eliminated gender bias, the impact of emotion was not detected due to poor recognition performance. These results demonstrate that audio LLMs are susceptible to making clinical decisions based on a patient's voice characteristics rather than medical evidence, a flaw that risks perpetuating healthcare disparities. We conclude that bias-aware architectures are essential and urgently needed before the clinical deployment of these models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。