arXiv:2608.18388cs.CV2026-08

用黎曼流匹配实现无需标注的动态4D场景重建

Depth Anything V4: Dynamic 4D Scene Reconstruction via Riemannian Flow Matching on 4D Gaussian Splatting

  • 在非欧空间对4D高斯点云参数建模,保证中间状态有效
  • 相比基线提升0.044的F-score,验证方法有效性
  • 适合做无标注动态场景重建的研究者与工程师

我们提出深度任意第四版(DAV4),一种从单目视频实现动态4D场景重建的框架。核心贡献是将黎曼流匹配(RFM)应用于4D高斯点云参数,在尺度、旋转、透明度等非欧流形上定义概率路径,确保所有中间状态合法。通过控制实验分离了RFM与测试时优化(TTO)及预训练的影响:相同数据、架构和TTO下,确定性MLP基线的F-score为0.762,而使用RFM达到0.806,净提升0.044即为RFM的独立贡献。我们提供修正后的计算成本分析:预训练耗时360 GPU小时,可支持大规模部署(超10,000场景)。不确定性通过负高斯对数似然与期望校准误差量化。DAV4在动态重建与新视角合成上超越此前深度任意模型及逐场景4D-GS方法,且训练过程中不依赖人工标注的深度标签。

原文摘要 · Abstract (English)

We present Depth Anything V4 (DAV4), a framework for dynamic 4D scene reconstruction from monocular video. Our key contribution is the application of Riemannian Flow Matching (RFM) to 4D Gaussian Splatting parameters, defining probability paths directly on non-Euclidean manifolds (scale, rotation, opacity), ensuring all intermediate states are valid. Through controlled experiments, we isolate RFM's contribution from test-time optimization (TTO) and pre-training. A deterministic MLP baseline with the same data, architecture, and TTO achieves F-score 0.762; RFM achieves 0.806 - the +0.044 gain is RFM's isolated contribution. We provide corrected computational cost analysis: pre-training is 360 GPU-hours, amortizing for large-scale deployment (over 10,000 scenes). Uncertainty is quantified via Negative Gaussian Log-Likelihood and Expected Calibration Error. DAV4 outperforms prior Depth Anything models and per-scene 4D-GS on dynamic reconstruction and novel-view synthesis, while using no human-annotated depth labels as training losses.

4D重建黎曼流匹配高斯溅射动态场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。