arXiv:2606.18558cs.CV2026-06被引 4

用语言指令预测3D点轨迹,提升机器人和生成模型的运动理解能力

MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction

论文配图:MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction
图 1 · 摘自论文原文
  • 基于语言指令与3D点轨迹,实现目标引导的运动预测
  • 在111类物体、61种运动上超越现有基线,准确率显著提升
  • 适用于机器人操作与视频生成,可迁移性强

运动预测是视觉智能的核心:智能体需预判物体运动以规划行为、推理物理交互并合成真实未来。本文认为,世界坐标系下的3D点具有类别无关、视角稳定、紧凑且直接可用的优势,提出目标条件下的3D点运动预测任务:给定短时视觉历史、目标物体上的3D查询点集及语言目标描述,模型预测每个点的未来3D轨迹。我们构建了完整研究体系:(1) MolmoMotion-1M,一个包含116万条未受控视频中标注的、带动作描述的3D点轨迹的大规模语料库;(2) PointMotionBench,一个涵盖111类物体和61种运动类型的人工验证基准;(3) MolmoMotion,一种支持自回归坐标预测与基于流匹配的轨迹生成的通用运动预测模型。MolmoMotion能准确预测多种运动模式,显著优于现有运动预测基线。最后,我们证明该3D运动先验在下游任务中具有良好迁移性:提升机器人操作训练效率与泛化能力,其预测轨迹还能为生成模型提供有效运动引导,合成更真实的物体运动视频。

原文摘要 · Abstract (English)

Motion forecasting is central to visual intelligence: agents must anticipate how objects will move in order to plan actions, reason about physical interactions, and synthesize realistic futures. We argue that 3D points in world coordinates provide a general representation that is class-agnostic, view-stable, compact, and directly useful for downstream tasks. We formalize the task of goal-conditioned 3D point motion forecasting: given a short visual history, a set of 3D query points on an object of interest, and a language description of the intended goal, the model predicts the future 3D trajectory of each point. We introduce a full stack to study this task at scale: (1) MolmoMotion-1M is a large corpus of action-described, object-grounded 3D point trajectories annotated from 1.16M unconstrained videos; (2) PointMotionBench is a human-verified benchmark spanning 111 object categories and 61 motion types; and (3) MolmoMotion is a general motion forecasting model that supports both autoregressive coordinate prediction and flow-matching-based trajectory generation. MolmoMotion accurately predicts diverse motion patterns with different language instructions, and significantly outperforms existing motion prediction baselines on PointMotionBench. Finally, we show that the learned 3D motion prior transfers well to downstream applications: it improves training efficiency and generalization for robot manipulation, and its predicted trajectories provide effective motion guidance for generative models to synthesize videos with more realistic object motion.

运动预测3D轨迹语言指令机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。