arXiv:2608.20331cs.CLcs.AI2026-08

让AI更懂病人:用检查清单+事实验证,生成准确易懂的医报解读。

G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation

论文配图:G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation
图 1 · 摘自论文原文
  • 用多源检索验证医学断言,结合用户上下文定制检查清单
  • 在真实数据集上,关键断言准确率提升17%,检查清单召回率达89%
  • 适合医疗AI落地、患者沟通优化与临床辅助系统研发者

个性化医疗报告解读对患者日益重要。现有视觉-语言任务难以兼顾事实准确性与患者沟通需求。为此,我们提出面向患者的医疗报告解释(PMRI)新任务,要求模型基于用户提问和对话历史,生成准确且易懂的解释。该任务需同时满足可验证性与情境适应性,传统监督微调与整体强化学习难以协同优化。为此,我们提出G-CARL框架:结合多源检索进行原子断言验证,利用上下文感知的实例化加权检查清单确保覆盖度,提供结构化监督而不限制表达多样性。我们构建了真实世界的MMedReport基准及三维度评估协议。实验表明,G-CARL在整体质量、断言级精确度和检查清单召回率上均优于现有后训练基线;临床医生偏好评测进一步证实其解释更准确、更贴合患者需求。

原文摘要 · Abstract (English)

Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressing this need requires both evidence-grounded medical factuality and context-dependent patient communication, yet existing medical vision-language tasks do not adequately capture these dual requirements. To bridge this gap, we introduce Patient-oriented Medical Report Interpretation (PMRI), a novel open-ended multimodal generation task that requires models to explain medical reports in accurate and accessible language based on a user's query and dialogue history. These two objectives differ fundamentally in their verifiability, yet remain tightly coupled, making them difficult to optimize jointly under conventional supervised fine-tuning and holistic reinforcement learning paradigms. To address this challenge, we propose G-CARL, a grounded, checklist-aligned reinforcement learning framework that combines multi-source retrieval for atomic claim verification with context-aware, instance-specific weighted checklists for response coverage, providing structured supervision for factuality, user-demand satisfaction, and expression quality without constraining response diversity. We further construct MMedReport, a real-world PMRI benchmark, along with a clinician-designed three-dimensional evaluation protocol. Extensive experiments demonstrate that G-CARL consistently outperforms existing post-training baselines in overall quality, claim-level precision, and checklist recall. Pairwise preference evaluation by clinicians further confirms that G-CARL produces interpretations that are more accurate and better aligned with patient needs.

医疗AI多模态生成强化学习患者沟通

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。