用点轨迹生成预测物体运动,比传统方法更准更多样。
What Happens Next? Anticipating Future Motion by Generating Point Trajectories
- 直接生成密集点轨迹,跳过像素重建环节
- 在模拟数据上准确率超现有方法,真实场景也表现良好
- 适合机器人规划、物理推理等需要运动预测的场景
我们研究从单张图像预测物体未来运动的问题,即不依赖速度或受力等参数,仅通过视觉信息推断物体可能的运动路径。将该任务建模为条件生成密集轨迹网格,采用接近现代视频生成器架构的模型,但输出为运动轨迹而非像素。该方法能捕捉全局动态与不确定性,相比以往回归器和生成器,预测更准确且多样性更高。我们在模拟数据上进行了广泛评估,验证了其在机器人等下游任务中的有效性,并在真实世界的直觉物理数据集上展现出良好精度。尽管当前先进的视频生成模型常被视为世界模型,但我们发现它们在从单图预测运动方面表现不佳,即使在落块或机械交互等简单物理场景中,即便经过微调也难以胜任。这表明其性能受限于生成像素的冗余开销,而非直接建模运动本身。
原文摘要 · Abstract (English)
We consider the problem of forecasting motion from a single image, i.e., predicting how objects in the world are likely to move, without the ability to observe other parameters such as the object velocities or the forces applied to them. We formulate this task as conditional generation of dense trajectory grids with a model that closely follows the architecture of modern video generators but outputs motion trajectories instead of pixels. This approach captures scene-wide dynamics and uncertainty, yielding more accurate and diverse predictions than prior regressors and generators. We extensively evaluate our method on simulated data, demonstrate its effectiveness on downstream applications such as robotics, and show promising accuracy on real-world intuitive physics datasets. Although recent state-of-the-art video generators are often regarded as world models, we show that they struggle with forecasting motion from a single image, even in simple physical scenarios such as falling blocks or mechanical object interactions, despite fine-tuning on such data. We show that this limitation arises from the overhead of generating pixels rather than directly modeling motion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。