arXiv:2507.21794cs.CV2025-07中稿 · MICCAI-W 2025被引 1

用结构化报告增强胸部X光图像理解,提升医疗多模态模型泛化能力。

Distribution-Based Masked Medical Vision-Language Model Using Structured Reports

  • 基于LLM生成的结构化报告,分三段注入疾病定义、关键区域和临床结论。
  • 通过建模跨模态与模态内不确定性,提升对医学数据模糊性的捕捉能力。
  • 适用于需要精准临床语义理解的医疗影像分析任务,如病灶检测与诊断辅助。

医疗图像-文本预训练旨在对齐医学图像与临床相关文本,以提升下游任务性能。然而现有模型常因医学数据的变异性与模糊性,难以捕捉细微临床信息及不确定性。本文提出一种考虑不确定性的医疗图像-文本预训练模型,专注于胸部X光,利用大语言模型(LLM)生成的结构化报告增强图像数据。报告包含疾病定义、关键区域(appearance)描述,以及观察结果(observations)与诊断结论(verdicts),使模型预测具临床语义基础。通过建模跨模态与模态内不确定性,该框架有效捕捉医学图像与文本中的固有模糊性,获得更优表征,在多个下游任务中实现当前最优表现。

原文摘要 · Abstract (English)

Medical image-language pre-training aims to align medical images with clinically relevant text to improve model performance on various downstream tasks. However, existing models often struggle with the variability and ambiguity inherent in medical data, limiting their ability to capture nuanced clinical information and uncertainty. This work introduces an uncertainty-aware medical image-text pre-training model that enhances generalization capabilities in medical image analysis. Building on previous methods and focusing on Chest X-Rays, our approach utilizes structured text reports generated by a large language model (LLM) to augment image data with clinically relevant context. These reports begin with a definition of the disease, followed by the `appearance' section to highlight critical regions of interest, and finally `observations' and `verdicts' that ground model predictions in clinical semantics. By modeling both inter- and intra-modal uncertainty, our framework captures the inherent ambiguity in medical images and text, yielding improved representations and performance on downstream tasks. Our model demonstrates significant advances in medical image-text pre-training, obtaining state-of-the-art performance on multiple downstream tasks.

医疗多模态图像文本对齐不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。