arXiv:2606.19908cs.CV2026-06

用高斯过程先验提升内窥镜视频修复效果,支持带不确定性的帧插补。

Gaussian Process Prior Variational Autoencoder for Endoscopic Videos

论文配图:Gaussian Process Prior Variational Autoencoder for Endoscopic Videos
图 1 · 摘自论文原文
  • 引入高斯过程先验替代传统独立潜变量,利用时间连续性建模帧间关系。
  • 在C3VDv2数据集上,重建均方根误差降低最多26.1%,轨迹误差减少12.7%。
  • 可输出每帧不确定性估计,适合临床辅助诊断与导航系统部署。

内窥镜视频分析对胃肠道诊断和计算机辅助干预至关重要,但视频序列常受镜面反光、运动伪影和丢帧等瞬时干扰影响。这些干扰会分散临床注意力、降低图像可读性,并破坏三维重建与导航等下游任务。有效修复需依赖时间连续性,而非孤立处理单帧。本文提出高斯过程先验变分自编码器(GPVAE)框架,以时间高斯过程先验替代标准因子化潜变量先验,实现带不确定性感知的缺失帧插补。该框架结合内窥镜专用编码器(包括卷积式EndoVAE主干和GastroNet-5M预训练视觉变换器),并采用两种可扩展的高斯过程近似方法:层级先验近似(HPA)和稀疏精度近似(SPA)。镜面反光通过基于DUCKNet的掩码流程处理,被污染像素被排除在重建目标之外。在C3VDv2结肠镜数据集上,最优的GPVAE变体相较匹配的VAE基线平均降低21.9%的图像重建RMSE,最高达26.1%;下游轨迹重建的RMSE平均降低12.7%,训练时间每轮平均增加27.3%。此外,GP后验提供每帧的不确定性估计,反映时间支持度,为修复结果提供置信信号。

原文摘要 · Abstract (English)

Endoscopic video analysis is essential for gastrointestinal diagnosis and computer-assisted interventions, but video sequences are routinely degraded by specular reflections, motion artifacts, and missing frames. These transient corruptions can distract clinicians, reduce image interpretability, and disrupt downstream tasks such as 3D reconstruction and navigation. Effective restoration therefore requires methods that exploit temporal continuity rather than treating frames in isolation. We introduce a Gaussian Process Prior Variational Autoencoder (GPVAE) framework for endoscopic video restoration that replaces the standard factorized latent prior with a temporal Gaussian process prior, enabling interpolation of missing frames with uncertainty-aware reconstruction. The framework combines endoscopy-specific encoders, including a convolutional EndoVAE backbone and pretrained Vision Transformer encoders from GastroNet-5M, with two scalable GP approximations: Hierarchical Prior Approximation (HPA) and Sparse Precision Approximation (SPA). Specular reflections are handled using a DUCKNet-based masking pipeline that excludes corrupted pixels from the reconstruction objective. On the C3VDv2 colonoscopy dataset, the best GPVAE variants reduced image reconstruction RMSE by 21.9\% on average, and by up to 26.1\%, relative to matched VAE baselines. Downstream trajectory RMSE was reduced by 12.7\% on average across classical visual odometry and a pretrained PoseNet, at an average increase of 27.3\% in training time per epoch. Finally, the GP posterior provides per-frame uncertainty estimates that reflect temporal support and offer a confidence signal for restored frames.

视频修复高斯过程内窥镜不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。