arXiv:2601.05237cs.CV2026-01被引 4

从人眼视角视频预测物体未来3D运动轨迹

ObjectForesight: Predicting Future 3D Object Trajectories from Human Videos

  • 基于3D物体中心建模,直接从视频推断物体6自由度运动
  • 在200万+视频片段上训练,实现高几何一致性预测
  • 适合需要理解物体交互的机器人与AR应用

人类能轻松预判物体在互动中的可能运动——想象杯子被拿起、刀子切开或盖子被合上。我们旨在赋予计算系统类似能力,仅通过被动视觉观察就预测物体未来的合理运动。提出ObjectForesight,一种3D物体中心的动力学模型,可从短时第一人称视频序列中预测刚性物体的6-DoF姿态与轨迹。不同于传统在像素或隐空间操作的世界/动力学模型,ObjectForesight在物体层面显式表示三维世界,实现几何基础且时间连贯的预测,捕捉物体功能与运动路径。为规模化训练,利用最新分割、网格重建和3D姿态估计技术,构建包含200万以上短片段的伪真值3D物体轨迹数据集。大量实验表明,ObjectForesight在准确性、几何一致性及对未见物体与场景的泛化能力上均有显著提升,建立了一个可扩展的框架,直接从观测中学习物理合理的物体中心动力学模型。

原文摘要 · Abstract (English)

Humans can effortlessly anticipate how objects might move or change through interaction--imagining a cup being lifted, a knife slicing, or a lid being closed. We aim to endow computational systems with a similar ability to predict plausible future object motions directly from passive visual observation. We introduce ObjectForesight, a 3D object-centric dynamics model that predicts future 6-DoF poses and trajectories of rigid objects from short egocentric video sequences. Unlike conventional world or dynamics models that operate in pixel or latent space, ObjectForesight represents the world explicitly in 3D at the object level, enabling geometrically grounded and temporally coherent predictions that capture object affordances and trajectories. To train such a model at scale, we leverage recent advances in segmentation, mesh reconstruction, and 3D pose estimation to curate a dataset of 2 million plus short clips with pseudo-ground-truth 3D object trajectories. Through extensive experiments, we show that ObjectForesight achieves significant gains in accuracy, geometric consistency, and generalization to unseen objects and scenes, establishing a scalable framework for learning physically grounded, object-centric dynamics models directly from observation. objectforesight.github.io

3D预测物体轨迹动作理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。