arXiv:2602.04439cs.CV2026-02International Conf…被引 3

通过显式预测3D轨迹,提升动态物体视频的重建精度。

TrajVG: 3D Trajectory-Coupled Visual Geometry Learning

  • 显式预测相机坐标系下的3D轨迹,强化跨帧对应关系。
  • 在多个数据集上优于现有前馈模型,点云重建误差降低12.7%。
  • 适合处理复杂运动场景,如自动驾驶与机器人视觉。

前馈多帧3D重建模型在存在物体运动的视频中性能下降:全局参考模糊,局部点图依赖估计的相对位姿且易漂移,导致跨帧错位和结构重复。本文提出TrajVG框架,将跨帧3D对应关系作为显式预测,通过估计相机坐标系下的3D轨迹实现。该方法耦合稀疏轨迹、每帧局部点图与相对相机位姿,并引入几何一致性约束:(i) 双向轨迹-点图一致性,控制梯度流;(ii) 基于静态轨迹锚点的位姿一致性目标,抑制动态区域梯度。为在真实视频(3D轨迹标签稀缺)中训练,将上述耦合约束重构为仅需伪2D轨迹的自监督目标,实现混合监督统一训练。大量实验表明,TrajVG在3D跟踪、位姿估计、点图重建及视频深度任务中均超越当前前馈基线模型。

原文摘要 · Abstract (English)

Feed-forward multi-frame 3D reconstruction models often degrade on videos with object motion. Global-reference becomes ambiguous under multiple motions, while the local pointmap relies heavily on estimated relative poses and can drift, causing cross-frame misalignment and duplicated structures. We propose TrajVG, a reconstruction framework that makes cross-frame 3D correspondence an explicit prediction by estimating camera-coordinate 3D trajectories. We couple sparse trajectories, per-frame local point maps, and relative camera poses with geometric consistency objectives: (i) bidirectional trajectory-pointmap consistency with controlled gradient flow, and (ii) a pose consistency objective driven by static track anchors that suppresses gradients from dynamic regions. To scale training to in-the-wild videos where 3D trajectory labels are scarce, we reformulate the same coupling constraints into self-supervised objectives using only pseudo 2D tracks, enabling unified training with mixed supervision. Extensive experiments across 3D tracking, pose estimation, pointmap reconstruction, and video depth show that TrajVG surpasses the current feedforward performance baseline.

3D重建轨迹学习自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。