arXiv:2511.00924cs.CL2025-11中稿 · NeurIPS被引 2

评估大模型在医患沟通中的理解力与共情能力,发现其存在认知复杂度高和情感偏见问题。

The Biased Oracle: Assessing LLMs' Understandability and Empathy in Medical Diagnoses

  • 用可读性指标与大模型评分结合评估理解力
  • 模型生成内容复杂且对不同人群情感回应不均
  • 适合关注医疗AI公平性与沟通质量的研究者

大型语言模型(LLMs)在辅助临床医生进行诊断沟通方面展现出潜力,能生成面向患者的解释与指导。然而,其输出是否具备可理解性与共情能力仍不明确。本文在医学诊断场景下评估了两个领先大模型,通过可读性指标衡量理解力,并以大模型作为裁判的评分结果对比人类评估来衡量共情能力。结果显示,模型会根据患者的社会人口学特征和病情调整解释内容,但同时也生成过于复杂的表述,表现出有偏差的情感共情,导致患者获取信息的可及性不均衡。这些现象凸显了系统性校准的必要性,以确保医患沟通的公平性。代码与数据已公开:https://github.com/Jeffateth/Biased_Oracle

原文摘要 · Abstract (English)

Large language models (LLMs) show promise for supporting clinicians in diagnostic communication by generating explanations and guidance for patients. Yet their ability to produce outputs that are both understandable and empathetic remains uncertain. We evaluate two leading LLMs on medical diagnostic scenarios, assessing understandability using readability metrics as a proxy and empathy through LLM-as-a-Judge ratings compared to human evaluations. The results indicate that LLMs adapt explanations to socio-demographic variables and patient conditions. However, they also generate overly complex content and display biased affective empathy, leading to uneven accessibility and support. These patterns underscore the need for systematic calibration to ensure equitable patient communication. The code and data are released: https://github.com/Jeffateth/Biased_Oracle

大模型医疗AI共情评估公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。