用1秒重建极稀疏视角下的3D CT,效果超当前方法1.5dB
X-LRM: X-ray Large Reconstruction Model for Extremely Sparse-View Computed Tomography Recovery in One Second
- 用X-former和X-triplane构建端到端3D重建框架,支持任意数量输入视图
- 在少于10个视角下实现27倍加速,比顶尖方法提升1.5dB
- 自建16K规模数据集,适合医学影像低剂量成像与快速重建场景
稀疏视角3D CT重建旨在从有限的2D X射线投影中恢复体数据。现有前馈方法受限于大规模训练数据稀缺及缺乏直接一致的3D表示。本文提出用于极端稀疏视角(<10视图)CT重建的X射线大重建模型(X-LRM)。X-LRM包含两个核心组件:基于MLP图像分词器与Transformer编码器的X-former,以及将输出特征上采样为隐式神经场表示的X-triplane。为支持训练,我们构建了包含超过16,000个体-投影对的大型数据集Torso-16K,覆盖多种躯干器官。大量实验表明,X-LRM相比当前最优方法提升1.5 dB,速度提高27倍,且具备更强灵活性。肺部分割任务评估进一步验证其实际应用价值。代码与数据集将在https://github.com/Richard-Guofeng-Zhang/X-LRM公开。
原文摘要 · Abstract (English)
Sparse-view 3D CT reconstruction aims to recover volumetric structures from a limited number of 2D X-ray projections. Existing feedforward methods are constrained by the scarcity of large-scale training datasets and the absence of direct and consistent 3D representations. In this paper, we propose an X-ray Large Reconstruction Model (X-LRM) for extremely sparse-view ($<$10 views) CT reconstruction. X-LRM consists of two key components: X-former and X-triplane. X-former can handle an arbitrary number of input views using an MLP-based image tokenizer and a Transformer-based encoder. The output tokens are then upsampled into our X-triplane representation, which models the 3D radiodensity as an implicit neural field. To support the training of X-LRM, we introduce Torso-16K, a large-scale dataset comprising over 16K volume-projection pairs of various torso organs. Extensive experiments demonstrate that X-LRM outperforms the state-of-the-art method by 1.5 dB and achieves 27$\times$ faster speed with better flexibility. Furthermore, the evaluation of lung segmentation tasks also suggests the practical value of our approach. Our code and dataset will be released at https://github.com/Richard-Guofeng-Zhang/X-LRM
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。