arXiv:2606.01204cs.CLcs.AI2026-06被引 1

LLM看病建议竟因语言不同差异巨大,暴露隐性地域偏见

Implicit Geographic Inference in LLM Medical Triage: Language-Driven Disparities in Emergency Recommendations

  • 用同一症状在6种语言提问,发现紧急就医推荐率从0%到30%不等
  • 指定美国地址使非英语提示的急诊建议率飙升76.7个百分点
  • 偏差源于语言暗示的地理判断,非翻译质量导致,适合关注AI公平性者阅读

我们研究大语言模型是否仅因患者提问语言不同,对相同症状给出不同医疗分诊建议。使用Gemini 3.5 Flash,在六种语言(英语、西班牙语、中文、印地语、日语、阿拉伯语)中测试神经科症状组合(持续头痛、视力模糊、恶心),每组执行30次(共450次API调用)。结果显示,尽管严重程度评分均在7.7-8.0/10之间,紧急就诊推荐率在日语和印地语中为0%,英语和阿拉伯语中达30%。添加“患者位于美国”的句子,使非英语提示的急诊建议率最高提升76.7个百分点;反之,英语提示加“东京”信息则将急诊率从30%降至6.7%。回译控制实验(日语转英语)结果与英语基线接近,证明差异并非由翻译质量引起,而是源于输入语言引发的隐性地理推断。我们公开全部数据集、实验代码与结果。

原文摘要 · Abstract (English)

We investigate whether large language models produce different medical triage recommendations for identical symptoms based solely on the language of the patient prompt. Using Gemini 3.5 Flash, we evaluate a neurological symptom profile (persistent headache, blurred vision, nausea) across six languages (English, Spanish, Chinese, Hindi, Japanese, Arabic) with 30 runs per condition (n=450 total API calls). We find that the model recommends emergency room visits at rates ranging from 0% (Japanese, Hindi) to 30% (English, Arabic), despite assigning nearly identical severity scores (7.7-8.0/10) across all languages. Adding a single sentence specifying the patient's US location increases ER recommendations by up to 76.7 percentage points for non-English prompts, while the reverse anchor (English prompt with a Tokyo location) reduces the ER rate from 30% to 6.7%. A back-translation control (Japanese to English) produces ER rates comparable to the English baseline, confirming that the disparity is not caused by translation quality but by implicit geographic inference from the input language. We release the complete dataset, experiment code, and results.

大模型偏见医疗AI语言差异地理推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。