arXiv:2602.23833eess.IVcs.CV2026-02中稿 · ance at MICCAI 202…

融合图像与元数据,提升医学影像序列分类的鲁棒性。

Revisiting Integration of Image and Metadata for DICOM Series Classification: Cross-Attention and Dictionary Learning

  • 用双向交叉注意力融合图像与元数据,显式建模跨模态交互。
  • 通过可学习词典和值条件调制处理缺失元数据,无需补全。
  • 支持不同长度和尺寸的序列,适用于真实医疗数据场景。

自动化识别DICOM影像序列对大规模医学图像分析、质量控制、协议标准化及可靠下游处理至关重要。然而,由于切片内容异质、序列长度可变,以及元数据完全缺失、不完整或不一致,分类任务仍具挑战。本文提出一种端到端多模态框架,联合建模图像内容与采集元数据,并显式应对上述问题:(i) 图像与元数据通过模态感知模块编码,以双向交叉注意力机制融合;(ii) 元数据由基于可学习特征字典的稀疏、缺失感知编码器处理,结合值条件调制,无需任何插补操作;(iii) 通过2.5D视觉编码器和等距采样切片上的注意力机制,处理序列长度与图像维度差异。在公开的Duke Liver MRI数据集和大型多中心内部队列上评估,涵盖域内性能与域外泛化能力。所有设置下,所提方法均显著优于仅图像、仅元数据及2D/3D多模态基线。结果表明,显式建模元数据稀疏性与跨模态交互能有效提升分类鲁棒性。

原文摘要 · Abstract (English)

Automated identification of DICOM image series is essential for large-scale medical image analysis, quality control, protocol harmonization, and reliable downstream processing. However, DICOM series classification remains challenging due to heterogeneous slice content, variable series length, and entirely missing, incomplete or inconsistent DICOM metadata. We propose an end-to-end multimodal framework for DICOM series classification that jointly models image content and acquisition metadata while explicitly accounting for all these challenges. (i) Images and metadata are encoded with modality-aware modules and fused using a bi-directional cross-modal attention mechanism. (ii) Metadata is processed by a sparse, missingness-aware encoder based on learnable feature dictionaries and value-conditioned modulation. By design, the approach does not require any form of imputation. (iii) Variability in series length and image data dimensions is handled via a 2.5D visual encoder and attention operating on equidistantly sampled slices. We evaluate the proposed approach on the publicly available Duke Liver MRI dataset and a large multi-institutional in-house cohort, assessing both in-domain performance and out-of-domain generalization. Across all evaluation settings, the proposed method consistently outperforms relevant image only, metadata-only and multimodal 2D/3D baselines. The results demonstrate that explicitly modeling metadata sparsity and cross-modal interactions improves robustness for DICOM series classification.

医学影像多模态元数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。