arXiv:2510.12818cs.CLcs.AI2025-10被引 5

通过改变患者代词测试大模型推理稳定性,发现即使诊断一致仍存在偏见。

MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning

  • 仅替换患者代词,保持症状不变,构建反事实测试集
  • 模型推理在风险因素和指南引用上出现局部差异,平均相似度超0.80
  • 适合医疗AI伦理审查与模型公平性评估

大型语言模型(LLMs)在临床决策支持中应用日益广泛,但细微的种族线索可能影响其推理过程。现有研究记录了不同患者群体间的输出差异,但对内部推理在受控人口特征变化下的演变知之甚少。本文提出MEDEQUALQA,一个反事实基准测试,仅扰动患者代词(he/him, she/her, they/them),同时固定关键症状与条件(CSCs)。每个临床案例被扩展为单个CSC消融,生成三个约23,000项的并行数据集(共69,000项)。我们评估GPT-4.1模型,并通过语义文本相似度(STS)衡量推理轨迹的稳定性。结果显示整体相似度高(均值STS >0.80),但在提及的风险因素、指南锚点及排序上存在持续的局部偏差,即使最终诊断未变。错误分析揭示若干推理路径转变的案例,突显可能引发不平等诊疗的临床偏见位置。MEDEQUALQA为医疗AI推理稳定性的审计提供了受控诊断场景。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed in clinical decision support, yet subtle demographic cues can influence their reasoning. Prior work has documented disparities in outputs across patient groups, but little is known about how internal reasoning shifts under controlled demographic changes. We introduce MEDEQUALQA, a counterfactual benchmark that perturbs only patient pronouns (he/him, she/her, they/them) while holding critical symptoms and conditions (CSCs) constant. Each clinical vignette is expanded into single-CSC ablations, producing three parallel datasets of approximately 23,000 items each (69,000 total). We evaluate a GPT-4.1 model and compute Semantic Textual Similarity (STS) between reasoning traces to measure stability across pronoun variants. Our results show overall high similarity (mean STS >0.80), but reveal consistent localized divergences in cited risk factors, guideline anchors, and differential ordering, even when final diagnoses remain unchanged. Our error analysis highlights certain cases in which the reasoning shifts, underscoring clinically relevant bias loci that may cascade into inequitable care. MEDEQUALQA offers a controlled diagnostic setting for auditing reasoning stability in medical AI.

医疗AI模型偏见反事实推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。