arXiv:2502.11655cs.CV2025-02被引 3

用文字指导物体中心视频预测,提升机器人场景理解能力

TextOCVP: Object-Centric Video Prediction with Language Guidance

  • 基于文本条件的Transformer模型,解析物体状态并预测未来
  • 在两个数据集上超越基线,实现更精准可控的视频生成
  • 结构化表示增强鲁棒性与可解释性,适合机器人应用

理解与预测未来场景状态对自主智能体在复杂环境中规划和行动至关重要。物体中心模型通过结构化潜在空间,在建模物体动态与预测未来场景方面展现出潜力,但常难以扩展至复杂真实场景,且难以融入外部引导信息,限制了其在机器人领域的应用。为此,我们提出TextOCVP——一种由文本描述引导的物体中心视频预测模型。TextOCVP将观测场景解析为称为“槽位”(slots)的物体表征,并利用文本条件的Transformer预测器,联合建模物体动态与交互关系,实现对物体状态及视频帧的未来预测。该方法在保持结构化潜在空间的同时,引入文本引导,提升了预测精度与可控性。在两个数据集上的实验表明,TextOCVP优于多个视频预测基线。此外,结构化的物体中心表征展现出更强的新型场景配置鲁棒性,以及更好的可控性与可解释性,使预测结果更精确、更易理解。视频与代码已公开于 https://play-slot.github.io/TextOCVP。

原文摘要 · Abstract (English)

Understanding and forecasting future scene states is critical for autonomous agents to plan and act effectively in complex environments. Object-centric models, with structured latent spaces, have shown promise in modeling object dynamics and predicting future scene states, but often struggle to scale beyond simple synthetic datasets and to integrate external guidance, limiting their applicability in robotics. To address these limitations, we propose TextOCVP, an object-centric model for video prediction guided by textual descriptions. TextOCVP parses an observed scene into object representations, called slots, and utilizes a text-conditioned transformer predictor to forecast future object states and video frames. Our approach jointly models object dynamics and interactions while incorporating textual guidance, enabling accurate and controllable predictions. TextOCVP's structured latent space offers a more precise control of the forecasting process, outperforming several video prediction baselines on two datasets. Additionally, we show that structured object-centric representations provide superior robustness to novel scene configurations, as well as improved controllability and interpretability, enabling more precise and understandable predictions. Videos and code are available at https://play-slot.github.io/TextOCVP.

视频预测物体中心文本引导机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。