用扩散模型生成伪多视角监督,解决单目视频动态重建难题
ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular Inputs
- 用个性化扩散模型生成伪多视角图像作为监督信号
- 在DyCheck上视觉质量与几何一致性超越现有方法
- 适合需要高精度动态场景重建的研究者
动态新视角合成旨在从任意视角生成运动主体的逼真视图。当依赖单目视频时,结构与运动解耦困难且监督稀少,挑战极大。本文提出视频扩散感知重建框架ViDAR,利用个性化扩散模型生成伪多视角监督信号,训练高斯点阵表示。通过引入场景特定特征进行条件控制,ViDAR恢复精细外观细节,同时减轻单目模糊带来的伪影。为解决扩散生成的时空不一致问题,设计了扩散感知损失函数和相机位姿优化策略,使合成视图与真实场景几何对齐。在具有极端视角变化的挑战性基准DyCheck上的实验表明,ViDAR在视觉质量和几何一致性上均优于所有先进基线。进一步验证其在动态区域重建上的显著优势,并建立了新基准以评估运动丰富区域的重建性能。
原文摘要 · Abstract (English)
Dynamic Novel View Synthesis aims to generate photorealistic views of moving subjects from arbitrary viewpoints. This task is particularly challenging when relying on monocular video, where disentangling structure from motion is ill-posed and supervision is scarce. We introduce Video Diffusion-Aware Reconstruction (ViDAR), a novel 4D reconstruction framework that leverages personalised diffusion models to synthesise a pseudo multi-view supervision signal for training a Gaussian splatting representation. By conditioning on scene-specific features, ViDAR recovers fine-grained appearance details while mitigating artefacts introduced by monocular ambiguity. To address the spatio-temporal inconsistency of diffusion-based supervision, we propose a diffusion-aware loss function and a camera pose optimisation strategy that aligns synthetic views with the underlying scene geometry. Experiments on DyCheck, a challenging benchmark with extreme viewpoint variation, show that ViDAR outperforms all state-of-the-art baselines in visual quality and geometric consistency. We further highlight ViDAR's strong improvement over baselines on dynamic regions and provide a new benchmark to compare performance in reconstructing motion-rich parts of the scene. Project page: https://vidar-4d.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。