arXiv:2603.09743cs.CV2026-03被引 1

用语言增强视觉理解,提升视频步骤规划准确率

LAP: A Language-Aware Planning Model For Procedure Planning In Instructional Videos

  • 将视觉画面转为文本描述,用语言特征指导动作规划
  • 在三个基准上达到领先性能,长序列规划效果显著
  • 适合需要精准动作顺序的智能助手、机器人应用

步骤规划需预测从初始视觉状态到目标状态的动作序列。现有方法主要依赖视觉输入,但不同动作在视觉上可能相似,导致歧义。本文提出语言感知规划(LAP),利用语言在潜在空间中更具区分性的特性来缓解此问题。LAP通过微调视觉语言模型(VLM)将视觉观察转化为文本描述,并生成动作预测与文本嵌入。这些文本嵌入比视觉嵌入更具有区分性,被用于扩散模型中规划动作序列。在CrossTask、Coin和NIV三个步骤规划基准上评估,LAP在多个指标和时间跨度下均大幅超越现有方法,验证了语言感知规划的有效性。

原文摘要 · Abstract (English)

Procedure planning requires a model to predict a sequence of actions that transform a start visual observation into a goal in instructional videos. While most existing methods rely primarily on visual observations as input, they often struggle with the inherent ambiguity where different actions can appear visually similar. In this work, we argue that language descriptions offer a more distinctive representation in the latent space for procedure planning. We introduce Language-Aware Planning (LAP), a novel method that leverages the expressiveness of language to bridge visual observation and planning. LAP uses a finetuned Vision Language Model (VLM) to translate visual observations into text descriptions and to predict actions and extract text embeddings. These text embeddings are more distinctive than visual embeddings and are used in a diffusion model for planning action sequences. We evaluate LAP on three procedure planning benchmarks: CrossTask, Coin, and NIV. LAP achieves new state-of-the-art performance across multiple metrics and time horizons by large margin, demonstrating the significant advantage of language-aware planning.

视频生成动作规划视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。