文化信息会误导大模型,导致医疗问答准确率下降3-7个百分点。
Counterfactual Cultural Cues Reduce Medical QA Accuracy in LLMs: Identifier vs Context Effects
- 用反事实数据增强测试集,加入文化标识和上下文线索
- 多模型测试显示文化信息使准确率下降3-7个百分点
- 适合关注医疗AI公平性与偏见的开发者与研究者
构建可持续且公平的医疗体系需要语言模型在面对非决定性文化信息时仍保持临床诊断正确。我们设计了一个反事实基准,将150个MedQA题目扩展为1650个变体,通过插入文化相关(i)标识符、(ii)上下文线索或(iii)两者组合,针对原住民加拿大人、中东穆斯林、东南亚三类群体,以及长度匹配的中性对照组进行测试,由临床医生验证所有变体的正确答案不变。评估GPT-5.2、Llama-3.1-8B、DeepSeek-R1和MedGemma(4B/27B)在仅选项与简要解释提示下的表现。结果显示,文化线索显著影响准确率(Cochran's Q, p<10^-14),当标识符与上下文共现时降幅最大(选项提示下达3-7个百分点),而中性修改仅引起微小非系统性变化。通过人类验证的评分标准(κ=0.76)由大模型判断发现,超过一半的文化相关解释最终导向错误答案,表明文化关联推理与诊断失败直接相关。我们公开了提示与增强数据,以支持对文化诱发诊断错误的评估与缓解。
原文摘要 · Abstract (English)
Engineering sustainable and equitable healthcare requires medical language models that do not change clinically correct diagnoses when presented with non-decisive cultural information. We introduce a counterfactual benchmark that expands 150 MedQA test items into 1650 variants by inserting culture-related (i) identifier tokens, (ii) contextual cues, or (iii) their combination for three groups (Indigenous Canadian, Middle-Eastern Muslim, Southeast Asian), plus a length-matched neutral control, where a clinician verified that the gold answer remains invariant in all variants. We evaluate GPT-5.2, Llama-3.1-8B, DeepSeek-R1, and MedGemma (4B/27B) under option-only and brief-explanation prompting. Across models, cultural cues significantly affect accuracy (Cochran's Q, $p<10^-14$), with the largest degradation when identifier and context co-occur (up to 3-7 percentage points under option-only prompting), while neutral edits produce smaller, non-systematic changes. A human-validated rubric ($κ=0.76$) applied via an LLM-as-judge shows that more than half of culturally grounded explanations end in an incorrect answer, linking culture-referential reasoning to diagnostic failure. We release prompts and augmentations to support evaluation and mitigation of culturally induced diagnostic errors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。