构建首个真实心理诊疗场景下的公平性评估数据集,揭示模型对患者性别等信息的偏见。
Moving Beyond Medical Exams: A Clinician-Annotated Fairness Dataset of Real-World Tasks and Ambiguity in Mental Healthcare
- 由临床专家创建,覆盖五大精神科决策领域,去除无关人口信息
- 16个现成模型测试显示,性别等变量显著影响模型决策,存在明显偏见
- 含不确定性标注,适合评估模型在模糊场景下的公平性与一致性
当前医疗语言模型评测多依赖模拟考试题,未能反映真实临床复杂性。尤其在精神科领域,模型易受患者性别、年龄等非决策因素干扰。为此,我们构建了一个美国本土化、无模型辅助的专家标注数据集,涵盖治疗、诊断、记录、监测和分诊五类核心决策任务。所有题目均移除与决策无关的人口统计信息(如性别、种族),替换为可变变量,支持对男性、女性及非二元性别患者的系统性评估。针对存在多种合理答案的模糊任务,我们引入专家标注的偏好数据集,包含不确定性标记。通过评估16个通用模型和6个心理健康微调模型在任务准确率、人口特征影响及自由回答一致性上的表现,验证了该数据集的可用性,揭示了现有模型在真实临床情境中的公平性缺陷。
原文摘要 · Abstract (English)
Current medical language model (LM) benchmarks often over-simplify the complexities of day-to-day clinical practice tasks and instead rely on evaluating LMs on multiple-choice board exam questions. In psychiatry especially, these challenges are worsened by fairness and bias issues, since models can be swayed by patient demographics even when those factors should not influence clinical decisions. Thus, we present an expert-created and annotated dataset spanning five critical domains of decision-making in mental healthcare: treatment, diagnosis, documentation, monitoring, and triage. This U.S.-centric dataset - created without any LM assistance - is designed to capture the nuanced clinical reasoning and daily ambiguities mental health practitioners encounter, reflecting the inherent complexities of care delivery that are missing from existing datasets. Almost all base questions with five answer options each have had the decision-irrelevant demographic patient information removed and replaced with variables, e.g., for age or ethnicity, and are available for male, female, or non-binary-coded patients. This design enables systematic evaluations of model performance and bias by studying how demographic factors affect decision-making. For question categories dealing with ambiguity and multiple valid answer options, we create a preference dataset with uncertainties from the expert annotations. We outline a series of intended use cases and demonstrate the usability of our dataset by evaluating sixteen off-the-shelf and six (mental) health fine-tuned LMs on category-specific task accuracy, on the fairness impact of patient demographic information on decision-making, and how consistently free-form responses deviate from human-annotated samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。