arXiv:2608.19825cs.CVcs.CL2026-08

提升医学影像描述的临床可靠性,通过双阶段对齐优化生成更精准的诊断描述。

Towards Clinically Faithful Medical Image Captioning via Enhanced Vision-Language Alignment

论文配图:Towards Clinically Faithful Medical Image Captioning via Enhanced Vision-Language Alignment
图 1 · 摘自论文原文
  • 分离训练与推理阶段的视觉语言对齐,增强临床语义一致性。
  • 引入概念级辅助学习和强化学习奖励机制,显著提升临床对齐度。
  • 无需额外训练即可通过重排序优化输出,适合临床部署场景。

医学图像描述技术可加速早期诊断流程并提升AI诊断系统的可解释性。然而,由于灰度模态、细微解剖线索、专业医学术语及数据质量差异,实现临床可靠的描述仍具挑战。尽管大视觉语言模型取得进展,流畅输出并不等同于临床概念空间对齐。为此,我们提出一种框架,通过分离并增强训练时与推理时的对齐。构建基于BioMedCLIP和SigLIP2的单/双视觉编码器、Q-Former与LLaMA解码器的流水线,并引入辅助学习以预测UMLS概念/类型。推理时采用单嵌入重排序筛选最优描述;训练时则引入MedPAIR-SCST,结合临床相关奖励,引导生成分布向更一致的临床描述倾斜。实验表明,多编码器设计与概念级辅助学习有助于保留临床信息;推理时重排序提供无需训练的对齐改进方法;而MedPAIR-SCST通过强化学习直接优化生成分布,提升描述的临床一致性。结果表明,联合使用选择式对齐与强化学习对齐,可在数据受限下实现更可信的医学图像描述。

原文摘要 · Abstract (English)

Medical image captioning is a technique that accelerates early-stage diagnostic workflows and enhances the interpretability of medical diagnostic AI systems. However, unlike general image captioning, clinically reliable captioning remains challenging due to grayscale-based modalities, subtle anatomical cues, specialized medical phrasing, and variations in data quality. Despite recent advances in large vision-language models, fluent outputs do not necessarily guarantee sufficient alignment with clinical concept spaces or evaluation criteria. To address this issue, we propose a framework that strengthens clinical alignment by separating and enhancing training-time alignment and inference-time alignment. We build a medical image captioning pipeline that integrates single/dual vision encoders based on BioMedCLIP and SigLIP2, a Q-Former, and a LLaMA-based decoder, and examine the contribution of auxiliary learning for UMLS concept/type prediction. At inference, we apply single-embedding-based reranking to select the best caption among candidates, while at training we introduce MedPAIR-SCST, which combines clinically relevant rewards to shift the generative distribution toward improved clinical alignment. Our experiments show that complementary visual representations with a multi-encoder design and concept-level auxiliary learning help preserve clinically meaningful information. Furthermore, inference-time reranking provides a practical way to improve semantic and clinical alignment without additional training, whereas MedPAIR-SCST goes beyond selection by directly improving the model's distribution to generate more consistent and clinically grounded captions. These findings suggest that jointly leveraging selection-based alignment and reinforcement-learning-based alignment can promote more trustworthy medical image captioning even in data-constrained settings.

医学图像视觉语言对齐生成模型临床可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。