arXiv:2509.10746cs.CL2025-09被引 3

让医疗对话模型像人一样思考情绪,提升可信度。

RECAP: Transparent Inference-Time Emotion Alignment for Medical Dialogue Systems

  • 基于认知评估理论分步解析患者输入,不需重新训练。
  • 8B到120B参数模型均提升情感对齐,大模型提升更明显。
  • 透明推理过程获专家青睐,适合临床部署的AI系统。

医疗大模型常生成情感平淡或不可解释的回复,难以获得临床信任。我们提出RECAP(Reflect-Extract-Calibrate-Align-Produce)框架,基于认知评估理论,在推理阶段将患者输入分解为可审计的、符合评估理论的多个步骤,无需重训练。在多个基准和8B至120B参数的模型上,RECAP显著提升与人类判断的一致性,且提升幅度随模型规模增大而增加。中间输出分析显示,模型普遍低估社会支持等关系因素。盲评中,肿瘤科医师对RECAP生成回复的评分显著高于基线,胜率76%-88%,证明有原则的提示方法能有效提升医疗AI的情绪智能,同时满足临床部署所需的透明性。

原文摘要 · Abstract (English)

Large language models in healthcare often produce emotionally flat or opaque responses, failing to provide the transparent reasoning required for clinical trust. We present RECAP (Reflect-Extract-Calibrate-Align-Produce), an inference-time framework grounded in cognitive appraisal theory that decomposes patient input into auditable, appraisal-theoretic stages without retraining. Across multiple benchmarks and models from 8B to 120B parameters, RECAP improves alignment with human judgments, with gains inversely proportional to model scale. Intermediate outputs further reveal that models systematically underweight relational factors such as social support. In blinded evaluations, oncology fellows rated RECAP responses significantly higher than baselines with 76-88% win rates, demonstrating that principled prompting can enhance medical AI's emotional intelligence while maintaining the transparency required for clinical deployment.

医疗对话情绪对齐透明推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。