用医学专家知识训练模型,自动检查病历生成是否合理。
Process-Supervised Reward Models for Verifying Clinical Note Generation: A Scalable Approach Guided by Domain Expertise
- 基于医学流程定义步骤,用专家知识注入错误数据
- 在两个评估中准确率超98%和56.2%,优于现有模型
- 适合医疗生成、大模型验证等需要可信输出的场景
过程监督奖励模型(PRM)在数学与编程等有标准答案的领域表现优异,但在缺乏真值答案的临床病历生成任务中面临挑战。本文提出一种新框架,通过明确定义有意义的‘步骤’、注入基于领域知识的合理‘错误’,并利用大语言模型规模化生成过程监督数据,克服了以往限制。所构建的基于LLaMA-3.1 8B的PRM,在两项关键评估中均达到领先水平:(1) 区分标准病历与含错样本的准确率达98.8%;(2) 选择医生偏好的病历准确率达56.2%。研究还探讨了有效训练的关键因素,包括损失函数设计与数据筛选策略,并开展全面的医生阅读实验,识别出影响下游Best-of-N性能的预测因子。该工作为跨领域生成任务中PRM的潜力释放提供了重要启示。
原文摘要 · Abstract (English)
Process-supervised reward models (PRMs) excel at providing step-by-step verification for large language model (LLM) outputs in domains like mathematics and coding. However, their application to fields lacking ground-truth answers, such as clinical note generation, poses significant challenges. We introduce a novel framework for training PRMs to deliver step-level reward signals for LLM-generated clinical notes. By precisely defining meaningful "steps," injecting realistic "errors" informed by domain expertise, and leveraging LLMs to generate process supervision data at scale, we overcome previous limitations. Our PRM, built on LLaMA-3.1 8B, consistently outperforms proprietary reasoning and non-reasoning models, achieving state-of-the-art performance on two key evaluations: (1) distinguishing gold-standard from error-containing samples with 98.8% accuracy, and (2) selecting physician-preferred clinical notes with 56.2% accuracy. We investigate critical components for effective PRM training, including optimal loss functions and data selection strategies, and present a comprehensive physician reader study identifying predictors of downstream Best-of-N performance. Our study sheds light on unlocking the potential of PRMs for diverse generative tasks across domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。