提出可应对影像缺失的解耦对齐模型,提升放射科报告生成准确性。
DiA-gnostic VLVAE: Disentangled Alignment-Constrained Vision Language Variational AutoEncoder for Robust Radiology Reporting with Missing Modalities
- 用专家混合架构解耦图像与临床模态特征
- 在两个数据集上分别达0.266和0.134的BLEU@4得分
- 适合临床数据不完整场景下的报告生成任务
将医学影像与临床背景结合是生成准确且具临床解释性的放射科报告的关键。然而,现有自动化方法常依赖资源密集型大语言模型或静态知识图谱,在真实临床数据中面临两大挑战:(1) 模态缺失,如临床信息不全;(2) 特征纠缠,即模态特异与共享信息混合,导致融合效果差且产生临床不符的幻觉。为此,我们提出DiA-gnostic VLVAE,通过解耦对齐实现鲁棒放射科报告生成。该框架基于多专家(MoE)视觉-语言变分自编码器(VLVAE),分离共享与模态特异性特征,并通过约束优化目标强制潜空间正交与对齐,防止次优融合。随后采用轻量级LLaMA-X解码器高效生成报告。在IU X-Ray和MIMIC-CXR数据集上,该方法分别取得0.266和0.134的BLEU@4分数,显著优于当前最优模型。
原文摘要 · Abstract (English)
The integration of medical images with clinical context is essential for generating accurate and clinically interpretable radiology reports. However, current automated methods often rely on resource-heavy Large Language Models (LLMs) or static knowledge graphs and struggle with two fundamental challenges in real-world clinical data: (1) missing modalities, such as incomplete clinical context , and (2) feature entanglement, where mixed modality-specific and shared information leads to suboptimal fusion and clinically unfaithful hallucinated findings. To address these challenges, we propose the DiA-gnostic VLVAE, which achieves robust radiology reporting through Disentangled Alignment. Our framework is designed to be resilient to missing modalities by disentangling shared and modality-specific features using a Mixture-of-Experts (MoE) based Vision-Language Variational Autoencoder (VLVAE). A constrained optimization objective enforces orthogonality and alignment between these latent representations to prevent suboptimal fusion. A compact LLaMA-X decoder then uses these disentangled representations to generate reports efficiently. On the IU X-Ray and MIMIC-CXR datasets, DiA has achieved competetive BLEU@4 scores of 0.266 and 0.134, respectively. Experimental results show that the proposed method significantly outperforms state-of-the-art models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。