用事实奖励优化临床文本生成,减少幻觉并提升完整性。
Optimizing Long-Form Clinical Text Generation with Claim-Based Rewards
- 基于事实点的奖励机制直接优化生成质量。
- 相比基准,事实性、完整性和简洁性显著提升。
- 适合医疗AI开发与临床文档自动化场景。
利用大语言模型自动化临床记录需精准对齐完整性与事实依据等目标。我们提出一种融合评估与强化学习的框架,结合群组相对策略优化(GRPO)与DocLens——一个提供确定性、对话相关的事实级奖励的评估器。该方法无需训练独立奖励模型或依赖人工参考,直接优化事实依据与完整性。实验表明,该方法提升了临床笔记质量并降低训练成本,通过简单的奖励门控策略实现。独立的GPT-5定性评估进一步验证其优势:在事实性、完整性和简洁性方面更受偏好,遗漏与幻觉更少。由于基准较干净且基础模型已良好对齐,此提升可能为保守下限。该框架可扩展至真实场景,并支持自定义目标如指南遵循或计费偏好。
原文摘要 · Abstract (English)
Automating clinical documentation with large language models requires precise alignment with priorities such as completeness and factual grounding. We present an evaluation-integrated reinforcement learning framework for long-form clinical text generation that couples Group Relative Policy Optimization (GRPO) with DocLens, a claim-level evaluator that provides deterministic, dialogue-grounded rewards. Our method directly optimizes factual grounding and completeness without training a separate reward model or relying on human-authored references. Empirically, the approach improves clinical note quality and reduces training cost via a simple reward-gating strategy. An independent GPT-5 qualitative evaluation further supports these gains, showing higher preference for GRPO outputs in factuality, completeness, and brevity, with fewer omissions and hallucinations. Because the benchmarks are relatively clean and the base model already well aligned, these improvements likely represent a conservative lower bound. The framework is scalable to real-world settings and can incorporate custom objectives such as guideline adherence or billing preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。