arXiv:2509.02007cs.AI2025-09被引 2

提出多维度医疗公平性评估框架,揭示模型在真实临床场景中的隐蔽偏见。

mFARM: Towards Multi-Faceted Fairness Assessment based on HARMs in Clinical Decision Support

  • 构建包含5万+提示的双基准测试集,覆盖12种种族性别组合与3类临床情境。
  • 提出mFARM评分体系,量化分配、稳定性和潜在三类医疗偏差,发现模型在上下文简化时表现骤降。
  • 适合关注医疗AI公平性、临床部署风险的研究者与开发者使用。

大型语言模型(LLMs)在高风险医疗场景中的应用带来严峻的AI对齐挑战,因模型可能继承并放大社会偏见,导致显著差异。现有公平性评估方法因依赖简单指标而忽视医疗伤害的多维特性,甚至催生仅因临床无害而看似公平但可能不准确的模型。为此,我们贡献两点:其一,基于MIMIC-IV构建两个大规模受控基准(ED-Triage和Opioid Analgesic Recommendation),涵盖超过50,000个提示,包含十二种种族×性别组合及三类情境层级;其二,提出多指标框架mFARM,用于审计分配性、稳定性与潜在性三类偏差,并整合为mFARM分数。同时引入公平性-准确性平衡(FAB)分数以衡量二者权衡。我们对四个开源模型(Mistral-7B、BioMistral-7B、Qwen-2.5-7B、Bio-LLaMA3-8B)及其微调版本在量化与上下文变化下的表现进行了实证评估。结果表明,所提mFARM指标在多种设置下更有效捕捉细微偏见;多数模型在不同量化水平下保持稳定的mFARM得分,但在上下文缩减时显著退化。相关基准与代码已公开,推动医疗领域对齐AI研究。

原文摘要 · Abstract (English)

The deployment of Large Language Models (LLMs) in high-stakes medical settings poses a critical AI alignment challenge, as models can inherit and amplify societal biases, leading to significant disparities. Existing fairness evaluation methods fall short in these contexts as they typically use simplistic metrics that overlook the multi-dimensional nature of medical harms. This also promotes models that are fair only because they are clinically inert, defaulting to safe but potentially inaccurate outputs. To address this gap, our contributions are mainly two-fold: first, we construct two large-scale, controlled benchmarks (ED-Triage and Opioid Analgesic Recommendation) from MIMIC-IV, comprising over 50,000 prompts with twelve race x gender variants and three context tiers. Second, we propose a multi-metric framework - Multi-faceted Fairness Assessment based on hARMs ($mFARM$) to audit fairness for three distinct dimensions of disparity (Allocational, Stability, and Latent) and aggregate them into an $mFARM$ score. We also present an aggregated Fairness-Accuracy Balance (FAB) score to benchmark and observe trade-offs between fairness and prediction accuracy. We empirically evaluate four open-source LLMs (Mistral-7B, BioMistral-7B, Qwen-2.5-7B, Bio-LLaMA3-8B) and their finetuned versions under quantization and context variations. Our findings showcase that the proposed $mFARM$ metrics capture subtle biases more effectively under various settings. We find that most models maintain robust performance in terms of $mFARM$ score across varying levels of quantization but deteriorate significantly when the context is reduced. Our benchmarks and evaluation code are publicly released to enhance research in aligned AI for healthcare.

医疗AI公平性评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。