用扩散模型从单张X光片生成多角度高质新视图。
SV-DRR: High-Fidelity Novel View X-Ray Synthesis Using Diffusion Model
- 基于扩散变换器,实现视角可控的高质量图像生成。
- 生成分辨率更高,视角控制更精准,超越以往方法。
- 适合临床诊断、医学教育及数据扩充场景使用。
X光成像是一种快速且低成本的内脏结构可视化工具。多视角X光可提供互补信息,提升诊断、手术与教学效果,但多角度采集增加辐射暴露并复杂化临床流程。为此,我们提出一种新型视角条件扩散模型,仅需单视角输入即可合成多视角X光图像。不同于先前方法在视角范围、分辨率和图像质量上的局限,本方法采用扩散变换器(Diffusion Transformer)以保留精细细节,并引入弱到强训练策略,实现稳定高分辨率生成。实验表明,该方法生成图像分辨率更高,视角控制更优。这一能力在临床应用、医学教育及数据扩展方面具有重要意义,可生成多样且高质量的数据集用于训练与分析。代码已公开于 https://github.com/xiechun298/SV-DRR。
原文摘要 · Abstract (English)
X-ray imaging is a rapid and cost-effective tool for visualizing internal human anatomy. While multi-view X-ray imaging provides complementary information that enhances diagnosis, intervention, and education, acquiring images from multiple angles increases radiation exposure and complicates clinical workflows. To address these challenges, we propose a novel view-conditioned diffusion model for synthesizing multi-view X-ray images from a single view. Unlike prior methods, which are limited in angular range, resolution, and image quality, our approach leverages the Diffusion Transformer to preserve fine details and employs a weak-to-strong training strategy for stable high-resolution image generation. Experimental results demonstrate that our method generates higher-resolution outputs with improved control over viewing angles. This capability has significant implications not only for clinical applications but also for medical education and data extension, enabling the creation of diverse, high-quality datasets for training and analysis. Our code is available at https://github.com/xiechun298/SV-DRR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。