arXiv:2606.01419cs.CV2026-06

基于深度引导的多模型融合,提升足球场景新视角生成质量

DENSER: Depth-Guided Ensemble with Staged EFA-GS Reconstruction for Soccer Novel View Synthesis

论文配图:DENSER: Depth-Guided Ensemble with Staged EFA-GS Reconstruction for Soccer Novel View Synthesis
图 1 · 摘自论文原文
  • 用相机高度加权损失,优先优化地面直播视角
  • 引入单目深度监督,改善无纹理区域的几何精度
  • 三模型像素平均集成,通过训练时长和尺度控制差异

我们提出 DENSER,一种用于足球新视角合成的深度引导集成方法。DENSER 在 EFA-GS 基础上做出三项改进:(1) 基于相机高度的损失加权,优先优化地面广播视角;(2) 使用 Depth-Anything-V2 提供单目深度监督,规范无纹理区域的几何结构;(3) 采用三个模型的像素平均集成,各模型从同一基础检查点出发,通过不同训练长度与高斯尺度钳制实现差异。在五个保留挑战场景中,平均达到 PSNR 29.89 dB、SSIM 0.791、LPIPS 0.366。

原文摘要 · Abstract (English)

We propose DENSER, a Depth-guided ENSemble with Staged EFA-GS Reconstruction for soccer novel view synthesis. DENSER extends EFA-GS with three key contributions: (1) camera-height-based loss weighting that prioritises ground-level broadcast views, (2) monocular depth supervision from Depth-Anything-V2 to regularise geometry in textureless regions, and (3) a three-model pixel-average ensemble whose members diverge from a shared base checkpoint by varying training length and Gaussian scale clamping. On five held-out challenge scenes we achieve a mean PSNR of 29.89 dB, SSIM of 0.791, and LPIPS of 0.366.

新视角合成深度引导足球视频多模型集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。