用大模型从少量投影图重建高质量CT,突破传统方法容量与灵活性限制。
X-GRM: Large Gaussian Reconstruction Model for Sparse-view X-rays to Computed Tomography
- 基于可扩展Transformer架构,高效融合多视角稀疏投影信息。
- 提出新型体素高斯点阵表示(VoxGS),支持可微分投影渲染与高效体积提取。
- 在域内与域外输入上均实现高质量重建,适合医学影像重建研究者使用。
计算机断层扫描在临床流程中不可或缺,可非侵入式呈现内部解剖结构。现有CT重建方法受限于小规模模型架构和僵化的体数据表示。本文提出X-GRM(X射线高斯重建模型),一种大型前馈模型,用于从稀疏视角2D X射线投影重建3D CT体数据。X-GRM采用可扩展的Transformer架构编码稀疏视角输入,不同视角的令牌被高效整合;随后这些令牌解码为一种新型体数据表示——体素高斯点阵(VoxGS),支持高效的CT体提取与可微分的X射线渲染。该大容量模型与灵活体表示的结合,使模型能在多种测试输入下生成高质量重建结果,包括域内与域外的X射线投影。代码已开源:https://github.com/CUHK-AIM-Group/X-GRM。
原文摘要 · Abstract (English)
Computed Tomography serves as an indispensable tool in clinical workflows, providing non-invasive visualization of internal anatomical structures. Existing CT reconstruction works are limited to small-capacity model architecture and inflexible volume representation. In this work, we present X-GRM (X-ray Gaussian Reconstruction Model), a large feedforward model for reconstructing 3D CT volumes from sparse-view 2D X-ray projections. X-GRM employs a scalable transformer-based architecture to encode sparse-view X-ray inputs, where tokens from different views are integrated efficiently. Then, these tokens are decoded into a novel volume representation, named Voxel-based Gaussian Splatting (VoxGS), which enables efficient CT volume extraction and differentiable X-ray rendering. This combination of a high-capacity model and flexible volume representation, empowers our model to produce high-quality reconstructions from various testing inputs, including in-domain and out-domain X-ray projections. Our codes are available at: https://github.com/CUHK-AIM-Group/X-GRM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。