arXiv:2503.15454cs.CL2025-03被引 8

检测并缓解医疗问答系统中的种族性别偏见,提升公平性。

Bias Evaluation and Mitigation in Retrieval-Augmented Medical Question-Answering Systems

  • 通过敏感属性查询评估检索增强系统的偏差
  • 多数投票法显著提升准确率与公平性指标
  • 适合关注医疗AI公平性的研究者与开发者

基于检索增强生成的医学问答系统在临床决策支持中具有潜力,因其可融合外部知识,减少独立大模型固有的不准确性。然而,这些系统可能无意中传播或放大与种族、性别及社会经济因素等敏感人口属性相关的偏见。本研究系统评估了多个问答基准(包括MedQA、MedMCQA、MMLU和EquityMedQA)中医疗RAG流程内的性别、种族等人口学偏差,通过生成并分析针对人口特征敏感的查询,量化了检索一致性与答案正确性的差异。进一步实施并比较了多种偏差缓解策略,包括思维链推理、反事实过滤、对抗性提示优化和多数投票聚合。实验结果表明存在显著的人口学差异,其中多数投票聚合方法显著提升准确率与公平性指标。研究强调需采用显式的公平性感知检索方法与提示工程策略,以构建真正公平的医疗问答系统。

原文摘要 · Abstract (English)

Medical Question Answering systems based on Retrieval Augmented Generation is promising for clinical decision support because they can integrate external knowledge, thus reducing inaccuracies inherent in standalone large language models (LLMs). However, these systems may unintentionally propagate or amplify biases associated with sensitive demographic attributes like race, gender, and socioeconomic factors. This study systematically evaluates demographic biases within medical RAG pipelines across multiple QA benchmarks, including MedQA, MedMCQA, MMLU, and EquityMedQA. We quantify disparities in retrieval consistency and answer correctness by generating and analyzing queries sensitive to demographic variations. We further implement and compare several bias mitigation strategies to address identified biases, including Chain of Thought reasoning, Counterfactual filtering, Adversarial prompt refinement, and Majority Vote aggregation. Experimental results reveal significant demographic disparities, highlighting that Majority Vote aggregation notably improves accuracy and fairness metrics. Our findings underscore the critical need for explicitly fairness-aware retrieval methods and prompt engineering strategies to develop truly equitable medical QA systems.

医疗AI偏见检测RAG公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。