arXiv:2601.10716cs.CV2026-01被引 5

自监督动态场景下视图合成新方法,有效消除鬼影与伪影。

WildRayZer: Self-supervised Large View Synthesis in Dynamic Environments

  • 通过相机仅渲染静态结构,残差识别瞬态区域并生成伪运动掩码。
  • 在15000段真实动态序列上训练,显著提升跨视角背景补全质量。
  • 适合需要高鲁棒性视图合成的自动驾驶、VR等应用。

我们提出WildRayZer,一种自监督框架,用于动态环境中新型视图合成(NVS),该环境同时存在相机与物体运动。动态内容破坏了静态NVS模型依赖的多视角一致性,导致鬼影、幻觉几何和不稳定的位姿估计。WildRayZer通过分析-合成测试解决此问题:仅使用相机的静态渲染器解释刚性结构,其残差揭示瞬态区域。基于这些残差,构建伪运动掩码,蒸馏出运动估计器,并用其遮蔽输入标记和门控损失梯度,使监督聚焦于跨视角背景补全。为支持大规模训练与评估,我们构建了动态现实地产10K(D-RE10K)数据集,包含15000段随意捕获的动态序列,以及配套的稀疏视图瞬态感知基准D-RE10K-iPhone。实验表明,WildRayZer在单次前馈中持续优于基于优化与前馈基线,在瞬态区域去除与全帧视图合成质量上表现更优。

原文摘要 · Abstract (English)

We present WildRayZer, a self-supervised framework for novel view synthesis (NVS) in dynamic environments where both the camera and objects move. Dynamic content breaks the multi-view consistency that static NVS models rely on, leading to ghosting, hallucinated geometry, and unstable pose estimation. WildRayZer addresses this by performing an analysis-by-synthesis test: a camera-only static renderer explains rigid structure, and its residuals reveal transient regions. From these residuals, we construct pseudo motion masks, distill a motion estimator, and use it to mask input tokens and gate loss gradients so supervision focuses on cross-view background completion. To enable large-scale training and evaluation, we curate Dynamic RealEstate10K (D-RE10K), a real-world dataset of 15K casually captured dynamic sequences, and D-RE10K-iPhone, a paired transient and clean benchmark for sparse-view transient-aware NVS. Experiments show that WildRayZer consistently outperforms optimization-based and feed-forward baselines in both transient-region removal and full-frame NVS quality with a single feed-forward pass.

视图合成自监督动态场景扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。