用单次推理生成术后脊柱影像,保持解剖结构准确
ViT3Flow: A Test-Time Training Transformer MeanFlow for Postoperative Radiograph Synthesis in Scoliosis

- 将手术矫正建模为有限区间生成传输,测试时训练自适应调整
- 在ScoliSurg数据集上达到最佳图像质量与几何精度
- 适合脊柱侧弯手术规划,兼顾效率与解剖真实性
从术前放射影像预测术后脊柱形态可为脊柱侧弯手术规划提供重要支持,但因手术矫正导致显著空间变化而结构需忠实保留,仍具挑战。本文将此问题定义为术后脊柱放射影像合成,并构建了首个针对该任务的成对数据集ScoliSurg,包含632对全脊柱术前-术后放射影像,附带结构化形态信息。提出ViT³Flow,一种仅需单次前向传播(single-NFE)的条件均值流框架,用于高效术后影像合成。该方法将手术矫正建模为有限区间生成传输,以测试时训练的令牌混合器替代传统自注意力,实现针对每例病例解剖与畸形模式的样本特异性内适应。此外,脊柱形态提取代理从术前影像提取主弯区域与方向的分布,指导诊断路由区间交叉注意力(DRICA),从独立的术前令牌流中进行区间依赖的垂直、水平、联合及全局检索。该设计使演进中的术后表征在整个传输过程中持续融入空间对应解剖证据。在ScoliSurg上的大量实验表明,ViT³Flow在感知图像质量、解剖保真度和临床相关几何精度上优于对比方法,且仅需一次网络评估。结果凸显其在脊柱侧弯手术规划中高效且解剖精准的术后影像合成潜力。
原文摘要 · Abstract (English)
Predicting postoperative spinal morphology from preoperative radiographs could provide valuable support for scoliosis surgical planning, but remains challenging because surgical correction induces large spatial changes while anatomical structures must be faithfully retained. We formulate this problem as postoperative scoliosis radiograph synthesis and construct ScoliSurg, the first paired dataset for this task, comprising 632 preoperative--postoperative whole-spine radiograph pairs with structured morphology information. We further propose ViT$^{3}$Flow, a single-NFE conditional MeanFlow framework for efficient postoperative radiograph synthesis. ViT$^{3}$Flow models surgical correction as finite-interval generative transport and replaces conventional self-attention with test-time-training token mixers that perform sample-specific inner adaptation to the anatomy and deformity pattern of each case. In addition, a Spinal Morphology Extraction Agent extracts distributions of dominant-curve region and direction from the preoperative radiograph. These distributions guide Diagnosis-Routed Interval Cross-Attention (DRICA), which performs interval-dependent vertical, horizontal, joint, and global retrieval from a separate preoperative token stream. This design enables the evolving postoperative representation to incorporate spatially corresponding anatomical evidence throughout the transport process. Extensive experiments on ScoliSurg demonstrate that ViT$^{3}$Flow achieves the best performance among the compared methods in perceptual image quality, anatomical fidelity, and clinically relevant geometric accuracy, while requiring only a single network evaluation. These results highlight the potential of ViT$^{3}$Flow for efficient and anatomically faithful postoperative radiograph synthesis in scoliosis surgical planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。