测试大模型如何响应患者价值陈述,发现响应虽表面积极但实际改变有限。
The Value Sensitivity Gap: How Clinical Large Language Models Respond to Patient Preference Statements in Shared Decision-Making
- 用临床案例测试4类大模型对13种价值取向的反应
- 模型仅小幅调整推荐,方向一致性最高达1.0
- 适合关注临床AI伦理与决策透明性的研究者
大型语言模型(LLMs)正进入临床决策支持流程,但其对患者明确表达的价值偏好——共享决策的核心内容——的响应尚未被量化。我们基于98,759条去标识化医疗补助就诊记录构建临床案例,测试了四种LLM家族(GPT-5.2、Claude 4.5 Sonnet、Gemini 3 Pro、DeepSeek-R1)在两个临床领域中对13种价值条件的反应,共完成104次实验。各模型默认价值取向差异显著(激进度量表2.0至3.5,满分5分)。价值敏感性指数介于0.13至0.27之间,与患者陈述偏好方向一致率在0.625至1.0之间。所有模型在非对照试验中均承认患者价值观,但实际推荐调整幅度较小。在78次二期评估中,决策矩阵与自我报告缓解策略各自使方向一致性提升0.125。这些结果为临床AI治理框架中提议的价值披露标签提供了实证依据。
原文摘要 · Abstract (English)
Large language models (LLMs) are entering clinical workflows as decision support tools, yet how they respond to explicit patient value statements -- the core content of shared decision-making -- remains unmeasured. We conducted a factorial experiment using clinical vignettes derived from 98,759 de-identified Medicaid encounter notes. We tested four LLM families (GPT-5.2, Claude 4.5 Sonnet, Gemini 3 Pro, and DeepSeek-R1) across 13 value conditions in two clinical domains, yielding 104 trials. Default value orientations differed across model families (aggressiveness range 2.0 to 3.5 on a 1-to-5 scale). Value sensitivity indices ranged from 0.13 to 0.27, and directional concordance with patient-stated preferences ranged from 0.625 to 1.0. All models acknowledged patient values in 100% of non-control trials, yet actual recommendation shifting remained modest. Decision-matrix and VIM self-report mitigations each improved directional concordance by 0.125 in a 78-trial Phase 2 evaluation. These findings provide empirical data for populating value disclosure labels proposed by clinical AI governance frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。