arXiv:2607.10840cs.CV2026-07中稿 · ECCV被引 3

实现大视角变化下任意时间点的动态三维重建。

OmniX: Any-view and Any-time 4D Reconstruction via Feed-forward Trajectory Fields

论文配图:OmniX: Any-view and Any-time 4D Reconstruction via Feed-forward Trajectory Fields
图 1 · 摘自论文原文
  • 分离动态运动与静态几何,用动态令牌建模3D轨迹场。
  • 在80K场景、128万视频数据上达到顶尖轨迹预测性能。
  • 适合做动态场景重建与视频理解的研究者使用。

现有前馈式4D重建方法要么仅预测每帧静态点云,忽略前景运动;要么虽能估计点云轨迹,但受限于小范围相机运动,难以在大视角变化下融合多时序观测。为此,本文提出OmniX,一种前馈式4D重建框架,可从具有大相机运动的视频中为每个像素预测稠密3D点轨迹。OmniX将动态运动建模与静态几何预测解耦,采用紧凑的动态令牌表示运动,利用3D运动的稀疏低秩特性,高效生成跨图像所有像素的轨迹场并保留全局交互。为支持训练,我们构建了基于UE5的自动4D数据引擎,推出大规模数据集,包含80,000个场景和1.28M个多视角视频,具备完整几何标注。OmniX在稠密3D点轨迹预测与3D点追踪任务上达到当前最优表现,并在视频深度估计与相机位姿估计上展现出竞争力。

原文摘要 · Abstract (English)

Previous feed-forward 4D reconstruction methods either predict per-frame static point clouds, ignoring foreground motion, or estimate point cloud trajectories while being limited to small camera motions. This restricts their ability to aggregate observations over time and reconstruct complete dynamic scenes under large viewpoint changes. To address this limitation, we propose OmniX, a feed-forward 4D reconstruction framework that predicts dense 3D point trajectories for every pixel from videos with large camera motion. OmniX decouples dynamic motion modeling from static geometry prediction and represents motion using a compact set of dynamic tokens. By leveraging the sparse and low-rank structure of 3D motion, these tokens generate trajectory fields for all pixels across all images while efficiently preserving global interactions. To facilitate training, we further build an automatic UE5-based 4D data engine and introduce a large-scale dataset containing 80K scenes and 1.28M multi-view videos with full geometric annotations. OmniX achieves state-of-the-art performance on dense 3D point trajectory prediction and 3D point tracking, while also demonstrating competitive results on video depth estimation and camera pose estimation.

4D重建动态场景轨迹预测视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。