通过精准对齐临床发现与视觉信息,提升医学报告生成的准确性与可靠性。
Linguistically-Aligned and Visually-Grounded Preference Optimization for Clinically-Augmented Medical Report Generation

- 基于实体级临床诊断模块,分离临床事实与语言表达差异。
- 引入视觉上下文触发的文本偏好反转机制,实现细粒度跨模态对齐。
- 针对高不确定预测实体进行反事实修正,有效降低潜在错误风险。
尽管医学报告生成(MRG)取得显著进展,但事实性错误仍限制其可靠性。现有基于直接偏好优化(DPO)的方法通常将模型生成报告与真实报告直接配对,无意中将关键临床发现与无关语言特征混杂,且缺乏显式多模态对齐。为此,我们提出DPO-Clin框架,聚焦于临床发现与跨模态对齐的偏好优化。首先,引入实体级临床诊断(ECD)模块,实现精准的事实性诊断,指导生成语言对齐的报告偏好对,分离临床差异与语言变化。其次,设计检索增强的多模态DPO(M2DPO),在视觉上下文切换时强制文本偏好反转,实现细粒度跨模态对齐。第三,定位正确但高度不确定的预测实体,通过反事实修改构建针对性偏好数据,以缓解潜在风险。在两个公开胸片数据集(MIMIC-CXR、IU X-Ray)和一个自建内窥镜数据集上的实验表明,DPO-Clin显著优于基线SFT模型,在临床感知指标上表现更优,并超越现有DPO方法,在不同基线架构和多种医学影像模态间展现强泛化能力。
原文摘要 · Abstract (English)
Despite significant advances in Medical Report Generation (MRG), the reliability remains constrained by the prevalence of factual errors. While Direct Preference Optimization (DPO) has emerged as a promising post-training paradigm to enhance the performance of Supervised Fine-Tuned (SFT) MRG models, existing DPO-based MRG methods typically adopt a naive preference construction that directly pairs model-generated reports with ground truth reports. This strategy inadvertently entangles critical clinical findings with clinically irrelevant linguistic characteristics, and fundamentally lacks explicit vision-language alignment. To address these challenges, we propose DPO-Clin, a novel post-training framework that focuses preference optimization on clinical findings and cross-modal alignment. First, we introduce the Entity-level Clinical Diagnostic (ECD) module to perform a precise entity-level factual diagnosis. ECD guides the generation of linguistically-aligned report preference pairs, isolating clinical discrepancies from linguistic variations. Second, to achieve fine-grained cross-modal alignment, we develop M2DPO, a retrieval-augmented multi-modal DPO variant that enforces textual preference inversion triggered by visual context switches. Third, we locate correct yet highly uncertain predicted entities and apply counterfactual modifications to construct targeted preference data for latent risk mitigation, thereby further enhancing the model reliability. Extensive experiments on two public chest X-ray datasets (MIMIC-CXR and IU X-Ray) and an in-house endoscopy dataset demonstrate that DPO-Clin significantly improves the SFT baselines on clinical-aware metrics. Furthermore, it achieves superior performance over existing DPO-based MRG methods, exhibiting robust generalizability across distinct baseline architectures and diverse medical imaging modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。