arXiv:2506.10633cs.CV2025-06被引 1

让胸部X光扩散模型更懂报告描述的定位

Anatomy-Grounded Weakly Supervised Prompt Tuning for Chest X-ray Latent Diffusion Models

论文配图:Anatomy-Grounded Weakly Supervised Prompt Tuning for Chest X-ray Latent Diffusion Models
图 1 · 摘自论文原文
  • 用弱监督提示调优提升文本与影像的对齐能力
  • 在MS-CXR上达新纪录,对未知数据也表现稳健
  • 适合医疗影像与多模态对齐研究者参考

近年来,潜在扩散模型在自然图像的文本引导生成中表现出色。然而,在医学影像领域,尤其是胸部X光这一模态,文本到图像的潜在扩散模型仍处于探索阶段,主要受限于数据稀缺(如隐私问题)。本文首先证明,标准的文本条件潜在扩散模型未能将自由文本放射科报告中的临床信息与扫描图像对应区域有效对齐。为解决此问题,我们提出一种微调框架,通过弱监督提示调优提升预训练模型的多模态对齐能力,使其可高效用于下游任务如短语定位。该方法在标准基准数据集MS-CXR上达到新最优性能,并在分布外数据集VinDr-CXR上同样表现出强鲁棒性。代码将公开共享。

原文摘要 · Abstract (English)

Latent Diffusion Models have shown remarkable results in text-guided image synthesis in recent years. In the domain of natural (RGB) images, recent works have shown that such models can be adapted to various vision-language downstream tasks with little to no supervision involved. On the contrary, text-to-image Latent Diffusion Models remain relatively underexplored in the field of medical imaging, primarily due to limited data availability (e.g., due to privacy concerns). In this work, focusing on the chest X-ray modality, we first demonstrate that a standard text-conditioned Latent Diffusion Model has not learned to align clinically relevant information in free-text radiology reports with the corresponding areas of the given scan. Then, to alleviate this issue, we propose a fine-tuning framework to improve multi-modal alignment in a pre-trained model such that it can be efficiently repurposed for downstream tasks such as phrase grounding. Our method sets a new state-of-the-art on a standard benchmark dataset (MS-CXR), while also exhibiting robust performance on out-of-distribution data (VinDr-CXR). Our code will be made publicly available.

医学影像扩散模型提示调优多模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。