用轻量级模型精准检测医学报告幻觉,提升生成质量与安全性
Process Reward Models for Sentence-Level Verification of LVLM Radiology Reports
- 基于临床上下文建模句子级事实正确性,实现细粒度验证
- 在MIMIC-CXR上比强基线提升7.5% MCC和1.8% AUROC
- 无需内部参数,通用性强,适合实际部署与报告筛选
利用大视觉语言模型(LVLM)自动化生成放射科报告潜力巨大,但常出现临床关键性幻觉,带来严重风险。现有幻觉检测方法往往缺乏句级粒度或对不同生成器泛化能力不足。本文提出一种新型句级过程奖励模型(Process Reward Model, PRM),可基于临床上下文和前文内容预测每句话的事实正确性。在MIMIC-CXR数据集上使用弱监督标签微调后,0.5B参数的轻量级PRM显著优于现有验证方法:例如,在某一LVLM输出上,马修斯相关系数(MCC)相对提升7.5%,AUROC提升1.8%。该模型不依赖内部模型状态,展现出对未见过的LVLM的强大泛化能力。进一步实证表明,基于PRM分数过滤最差10%报告,可使F1-CheXbert得分提升4.5%;在新提出的加权最佳N选一策略中,临床指标(F1-CheXbert)相对提升7.4%,BERTScore提升0.6%。结果表明,轻量级、上下文感知的PRM为临床级LVLM提供了无需访问内部激活的通用安全层。
原文摘要 · Abstract (English)
Automating radiology report generation with Large Vision-Language Models (LVLMs) holds great potential, yet these models often produce clinically critical hallucinations, posing serious risks. Existing hallucination detection methods frequently lack the necessary sentence-level granularity or robust generalization across different LVLM generators. We introduce a novel approach: a sentence-level Process Reward Model (PRM) adapted for this vision-language task. Our PRM predicts the factual correctness of each generated sentence, conditioned on clinical context and preceding text. When fine-tuned on MIMIC-CXR with weakly-supervised labels, a lightweight 0.5B-parameter PRM outperforms existing verification techniques, demonstrating, for instance, relative improvements of 7.5% in Matthews Correlation Coefficient and 1.8% in AUROC over strong white-box baselines on outputs from one LVLM. Unlike methods reliant on internal model states, our PRM demonstrates strong generalization to an unseen LVLM. We further show its practical utility: PRM scores effectively filter low-quality reports, improving F1-CheXbert scores by 4.5% (when discarding the worst 10% of reports). Moreover, when guiding a novel weighted best-of-N selection process on the MIMIC-CXR test set, our PRM show relative improvements in clinical metrics of 7.4% for F1-CheXbert and 0.6% for BERTScore. These results demonstrate that a lightweight, context-aware PRM provides a model-agnostic safety layer for clinical LVLMs without access to internal activations
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。