arXiv:2602.14021cs.CV2026-02被引 6

用场景光流统一动态3D重建与跟踪,一次推理搞定几何与运动。

Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow

  • 以场景光流为核心,联合建模3D结构与相机/物体运动。
  • 单次前向传播完成几何与双向运动推断,无需显式位姿回归。
  • 在静态和动态数据上联合训练,4D重建与跟踪性能领先。

动态3D场景的重建与跟踪是计算机视觉的核心挑战。现有方法通常将几何与运动分离:静态多视角重建系统假设世界刚性,而动态跟踪框架依赖显式的自身运动估计或独立的对象运动模型。本文提出Flow4R,一种统一框架,将相对场景光流作为连接3D结构、相机自身运动与动态物体运动的核心表示。给定双视图输入,Flow4R使用共享的Vision Transformer预测一个紧凑的像素对齐属性集,包括3D点位置、场景光流、位姿权重和置信度图。该光流中心范式使得局部几何与双向运动可在单次前向传播中联合推断,无需显式位姿回归头或复杂束调整。通过在静态与动态数据集上联合训练,Flow4R在4D重建与跟踪基准上达到当前最优性能,验证了光流中心范式在时空场景理解中的强大能力。

原文摘要 · Abstract (English)

Reconstructing and tracking dynamic 3D scenes is a fundamental challenge in computer vision. Existing methods typically decouple geometry from motion: static multi-view reconstruction systems assume a rigid world, whereas dynamic tracking frameworks rely on explicit ego-motion estimation or separate object motion models. In this work, we propose Flow4R, a unified framework that treats relative scene flow as the central representation linking 3D structure, camera ego-motion, and dynamic object motion. Given a two-view input, Flow4R employs a shared Vision Transformer to predict a compact, pixel-aligned property set comprising 3D point positions, scene flow, pose weights, and confidence maps. This flow-centric formulation allows local geometry and bidirectional motion to be jointly inferred in a single feedforward pass, eliminating the need for explicit pose regression heads or complex bundle adjustment. By training jointly on static and dynamic datasets, Flow4R achieves state-of-the-art performance on 4D reconstruction and tracking benchmarks, demonstrating the power of the flow-centric formulation for spatiotemporal scene understanding.

4D重建场景光流统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。