arXiv:2511.14918cs.CV2025-11中稿 · CVPR被引 2

用CT生成2D胸片,让模型学会预测3D解剖结构。

X-WIN: Building Chest Radiograph World Model via Predictive Sensing

  • 从CT中学习3D结构,预测不同视角的2D胸片。
  • 在多种下游任务上超越现有模型,少样本微调效果好。
  • 适合医学影像研究者,提升疾病诊断能力。

胸部X光(CXR)是疾病诊断的重要医学成像技术,但作为二维投影图像,受限于结构重叠,难以捕捉三维解剖结构,导致表征学习和疾病诊断困难。为此,我们提出新型胸部X光世界模型X-WIN,通过在潜在空间中学习从胸部计算机断层扫描(CT)生成其二维投影,从而提取体积知识。核心思想是:具备内部化3D解剖结构知识的世界模型,可在三维空间中预测不同变换下的胸片。在投影预测中,引入亲和引导对比对齐损失,利用同一体积各投影间的相互相似性,捕获丰富的相关信息。为增强模型适应性,通过掩码图像建模将真实胸片纳入训练,并使用领域分类器促使真实与模拟胸片具有统计上相似的表示。全面实验表明,X-WIN在多种下游任务中,通过线性探测和少样本微调均优于现有基础模型。此外,X-WIN还具备重建3D CT体积的2D投影渲染能力。

原文摘要 · Abstract (English)

Chest X-ray radiography (CXR) is an essential medical imaging technique for disease diagnosis. However, as 2D projectional images, CXRs are limited by structural superposition and hence fail to capture 3D anatomies. This limitation makes representation learning and disease diagnosis challenging. To address this challenge, we propose a novel CXR world model named X-WIN, which distills volumetric knowledge from chest computed tomography (CT) by learning to predict its 2D projections in latent space. The core idea is that a world model with internalized knowledge of 3D anatomical structure can predict CXRs under various transformations in 3D space. During projection prediction, we introduce an affinity-guided contrastive alignment loss that leverages mutual similarities to capture rich, correlated information across projections from the same volume. To improve model adaptability, we incorporate real CXRs into training through masked image modeling and employ a domain classifier to encourage statistically similar representations for real and simulated CXRs. Comprehensive experiments show that X-WIN outperforms existing foundation models on diverse downstream tasks using linear probing and few-shot fine-tuning. X-WIN also demonstrates the ability to render 2D projections for reconstructing a 3D CT volume.

医学影像世界模型3D重建生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。