arXiv:2604.26232cs.CVcs.AI2026-04被引 1

让肠镜视频生成既可控又可解释,真实还原解剖结构。

DepthPilot: From Controllability to Interpretability in Colonoscopy Video Generation

论文配图:DepthPilot: From Controllability to Interpretability in Colonoscopy Video Generation
图 1 · 摘自论文原文
  • 用深度约束对扩散模型微调,保证解剖结构准确性。
  • 采用可学习样条去噪模块,捕捉复杂时空动态。
  • 生成视频临床可解释性强,适合医疗训练与导航。

可控医疗视频生成虽取得显著进展,但仍缺乏可解释性,即生成内容需符合物理先验与真实临床表现。为推动从可控性迈向可解释性,本文提出首个可解释的肠镜视频生成框架DepthPilot。该框架通过两种协同机制实现:一是设计先验分布对齐策略,通过参数高效微调将深度约束注入扩散主干网络,实现显式几何定位;二是引入自适应样条去噪模块,以可学习样条函数替代固定线性权重,增强在几何约束下的非线性建模能力。在三个公开数据集及内部临床数据上的广泛评估表明,DepthPilot能生成高度物理一致的视频,在所有基准上FID分数低于15,且在医生评估中排名第一,弥合了“视觉真实”与“临床可解释”之间的鸿沟。此外,其生成视频有望支持可靠3D重建,助力手术导航与盲区识别,为构建结直肠世界模型奠定基础。

原文摘要 · Abstract (English)

Controllable medical video generation has achieved remarkable progress, but it still lacks interpretability, which requires the alignment of generated contents with physical priors and faithful clinical manifestations. To push the boundaries from mere controllability to interpretability, we propose DepthPilot, the first interpretable framework for colonoscopy video generation. This work takes a step toward trustworthy generation through two synergistic paradigms. To achieve explicit geometric grounding, DepthPilot devises a prior distribution alignment strategy, injecting depth constraints into the diffusion backbone via parameter-efficient fine-tuning to ensure anatomical fidelity. To enhance intrinsic nonlinear modeling under these geometric constraints, DepthPilot employs an adaptive spline denoising module, replacing fixed linear weights with learnable spline functions to capture complex spatio-temporal dynamics. Extensive evaluations across three public datasets and in-house clinical data confirm DepthPilot's robust ability to produce physically consistent videos. It achieves FID scores below 15 across all benchmarks and ranks first in clinician assessments, bridging the gap between "visually realistic" and "clinically interpretable". Moreover, DepthPilot-generated videos are expected to enable reliable 3D reconstruction, facilitating surgical navigation and blind region identification, and serve as a foundation toward the colorectal world model.

医学视频可解释生成扩散模型肠镜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。