arXiv:2603.10494cs.CLcs.LG2026-03

用噪声验证反馈训练模型,让医疗摘要更真实且不缩水。

Coverage-Controlled Preference Mining from Noisy Claim Verification for Evidence-Grounded Generation

  • 将噪声验证结果转化为受控覆盖度的摘要偏好对
  • 临床摘要未支持率从10.7%降至1.9%,事实一致性提升明显
  • 适合医疗生成、需要高可信输出的研究者使用

证据扎根生成要求摘要中的主张必须有对应证据支持,但逐条验证器提供噪声反馈,可能奖励输出更少的模型。本文在临床简要住院病程摘要任务中研究该问题,要求输出基于患者电子病历(EHR)证据。提出VERI-DPO框架,将噪声声明验证转化为受控覆盖度的摘要级偏好。对每个证据窗口提示,采样多个候选摘要,分解为主张,逐项验证是否与患者证据一致,并仅当所选摘要具有更高整体验证支持度且保留相近可验证内容时形成偏好对。标准直接偏好优化将这些对提炼为单样本策略,避免推理时重排序。在患者独立的MIMIC-III-Ext-VeriFact-BHC测试集上,VERI-DPO使未支持率从10.7%降至1.9%(基于挖掘验证器),从11.6%降至6.4%(基于GPT-4o独立评判)。两名领域研究人员在100次盲评中,56次更偏好VERI-DPO而非基线模型,用于事实忠实性。在无模型适配的锁定零样本迁移测试(MIMIC-IV-Ext-BHC,1000例患者)中,未支持率下降,评分主张数量基本不变。多种子消融实验表明,验证器引导的配对构建是性能提升关键,覆盖控制与反退化机制防止了因输出变短或不可检而出现的虚假事实提升。

原文摘要 · Abstract (English)

Evidence-grounded generation produces summaries whose claims should be supported by supplied evidence, but claim-level verifiers provide noisy feedback and can reward models that simply say less. We study this problem in clinical Brief Hospital Course summarization, where outputs must remain grounded in patient-specific EHR evidence. We introduce VERI-DPO, a preference-mining framework that converts noisy claim verification into coverage-controlled summary-level preferences. For each evidence-window prompt, VERI-DPO samples multiple candidate summaries, decomposes them into claims, verifies each claim against patient evidence, and forms a preference pair only when the chosen summary has better aggregate verifier-estimated support while retaining comparable verifiable content. Standard Direct Preference Optimization then distills these pairs into a single-sample policy, avoiding inference-time reranking. On patient-disjoint MIMIC-III-Ext-VeriFact-BHC test data, VERI-DPO reduces Not Supported rates from 10.7% to 1.9% under the mining verifier and from 11.6% to 6.4% under a separately prompted GPT-4o judge. In 100 blinded pairwise assessments by two domain researchers, VERI-DPO is preferred over the base model 56 times versus 18 times for factual faithfulness. In a locked zero-shot MIMIC-IV-Ext-BHC transfer test with 1,000 patients and no model adaptation, VERI-DPO lowers Not Supported rates with nearly unchanged scored-claim counts. Multi-seed ablations show that verifier-guided pair construction drives the gains, while coverage and anti-degeneration controls prevent apparent factuality improvements from coming from shorter or less checkable outputs.

医疗生成偏好学习证据扎根真实性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。