arXiv:2503.06637cs.CV2025-03中稿 · RO-MAN 2026被引 5

用视觉语言联合建模,让扩散模型更准规划操作步骤

CLAD: Constrained Latent Action Diffusion for Vision-Language Procedure Planning

  • 用VAE学习动作与观测的潜在表示作为约束
  • 在跨任务、硬币、NIV数据集上显著超越现有方法
  • 适合需要多模态交互规划的AI系统研究者

我们提出CLAD,一种用于教学视频中视觉-语言流程规划的受限潜在动作扩散模型。流程规划是在给定起始和目标状态视觉观察的基础上预测中间动作的挑战性任务。未来交互式AI系统还需能结合多模态输入(如视觉加语言描述)进行规划。为此,我们的方法利用变分自编码器(VAE)学习动作与观测的潜在表示,并将其作为约束融入扩散过程。该方法利用扩散模型潜在空间中已有的语义信息,通过潜在约束引导扩散模型生成更准确的动作。我们在CrossTask、Coin和NIV三个主流数据集上进行了广泛实验,结果表明该方法显著优于现有最先进方法。通过消融实验进一步证明,基于VAE学习的动作与观测表示在潜在空间中的集成是性能提升的关键。

原文摘要 · Abstract (English)

We propose CLAD, a Constrained Latent Action Diffusion model for vision-language procedure planning in instructional videos. Procedure planning is the challenging task of predicting intermediate actions given a visual observation of a start and a goal state. However, future interactive AI systems must also be able to plan procedures using multi-modal input, e.g., where visual observations are augmented with language descriptions. To tackle this vision-language procedure planning task, our method uses a Variational Autoencoder (VAE) to learn the latent representation of actions and observations as constraints and integrate them into the diffusion process. This approach exploits that the latent space of diffusion models already has semantics that can be used. We use the latent constraints to steer the diffusion model to better generate actions. We report extensive experiments on the popular CrossTask, Coin, and NIV datasets and show that our method outperforms state-of-the-art methods by a large margin. By evaluating ablated versions of our method, we further show that the proposed integration of the action and observation representations learnt in the VAE latent space is key to these performance improvements.

视觉语言流程规划扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。