用人类反馈提升锥光束CT转多排CT的图像质量
Human-Guided Shade Artifact Suppression in CBCT-to-MDCT Translation via Schrödinger Bridge with Conditional Diffusion
- 结合生成对抗网络与条件扩散模型,通过边界一致性约束保持解剖结构准确
- 仅需10步采样即达到优于以往方法的精度,多指标全面领先
- 无需奖励模型,通过人类二值反馈实现偏好学习,适合临床医生参与
我们提出一种基于薛定谔桥(Schrödinger Bridge, SB)框架的CBCT到MDCT图像转换新方法,融合生成对抗网络先验与人类引导的条件扩散模型。不同于传统GAN或扩散模型,该方法显式保证输入CBCT与伪目标之间的边界一致性,确保解剖结构保真度与感知可控性。通过分类器自由引导(CFG)引入二值人类反馈,有效引导生成过程向临床期望结果收敛。采用迭代优化与锦标赛式偏好选择机制,模型在不依赖奖励模型的情况下内化人类偏好。减影图像可视化显示,该方法能选择性抑制关键解剖区域的阴影伪影,同时保留细微结构细节。定量评估表明,在临床数据集上,本方法在RMSE、SSIM、LPIPS和Dice指标上均优于先前的GAN及微调类反馈方法,且仅需10次采样步骤。结果证明该框架在实时性与偏好对齐方面具有显著优势。
原文摘要 · Abstract (English)
We present a novel framework for CBCT-to-MDCT translation, grounded in the Schrodinger Bridge (SB) formulation, which integrates GAN-derived priors with human-guided conditional diffusion. Unlike conventional GANs or diffusion models, our approach explicitly enforces boundary consistency between CBCT inputs and pseudo targets, ensuring both anatomical fidelity and perceptual controllability. Binary human feedback is incorporated via classifier-free guidance (CFG), effectively steering the generative process toward clinically preferred outcomes. Through iterative refinement and tournament-based preference selection, the model internalizes human preferences without relying on a reward model. Subtraction image visualizations reveal that the proposed method selectively attenuates shade artifacts in key anatomical regions while preserving fine structural detail. Quantitative evaluations further demonstrate superior performance across RMSE, SSIM, LPIPS, and Dice metrics on clinical datasets -- outperforming prior GAN- and fine-tuning-based feedback methods -- while requiring only 10 sampling steps. These findings underscore the effectiveness and efficiency of our framework for real-time, preference-aligned medical image translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。