arXiv:2412.11621cs.CVcs.MM2024-12AAAI被引 4

用图文提示让大模型生成连贯的步骤计划,提升任务规划效果。

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting

  • 结合文本与视频生成能力,通过桥梁机制实现跨模态协同。
  • 在新数据集Daily-PP上优于单一模态基线,提升计划准确性和连贯性。
  • 适合需要图文联合规划的应用,如智能助手和自动化流程设计。

基于大语言模型(LLM)的智能体在程序化任务中展现出潜力,但文本与视频联合指令对用户协助的潜力仍未充分探索。为此,我们提出视觉引导的文-视频提示(VG-TVP)方法,一种由大语言模型驱动的多模态程序化规划(MPP)框架。该框架可根据高层目标生成连贯的文本与视频程序计划。主要挑战在于确保文本与视觉信息量、时间连贯性及计划准确性。VG-TVP利用大模型的零样本推理能力、视频转文本生成能力以及扩散模型的文本转视频生成能力。通过提出新型的描述融合(FoC)方法,以及文本到视频桥(T2V-B)和视频到文本桥(V2T-B),增强模态间交互,使大模型能指导视觉化文本计划与文本化视频计划的生成。为应对多模态程序化规划数据稀缺问题,我们构建了新的数据集Daily-Life Task Procedural Plans(Daily-PP)。通过全面实验与基准测试,评估人类偏好(包括文本与视觉信息量、时间连贯性及计划准确性)。结果表明,我们的VG-TVP方法在Daily-PP数据集上优于单模态基线。

原文摘要 · Abstract (English)

Large Language Model (LLM)-based agents have shown promise in procedural tasks, but the potential of multimodal instructions augmented by texts and videos to assist users remains under-explored. To address this gap, we propose the Visually Grounded Text-Video Prompting (VG-TVP) method which is a novel LLM-empowered Multimodal Procedural Planning (MPP) framework. It generates cohesive text and video procedural plans given a specified high-level objective. The main challenges are achieving textual and visual informativeness, temporal coherence, and accuracy in procedural plans. VG-TVP leverages the zero-shot reasoning capability of LLMs, the video-to-text generation ability of the video captioning models, and the text-to-video generation ability of diffusion models. VG-TVP improves the interaction between modalities by proposing a novel Fusion of Captioning (FoC) method and using Text-to-Video Bridge (T2V-B) and Video-to-Text Bridge (V2T-B). They allow LLMs to guide the generation of visually-grounded text plans and textual-grounded video plans. To address the scarcity of datasets suitable for MPP, we have curated a new dataset called Daily-Life Task Procedural Plans (Daily-PP). We conduct comprehensive experiments and benchmarks to evaluate human preferences (regarding textual and visual informativeness, temporal coherence, and plan accuracy). Our VG-TVP method outperforms unimodal baselines on the Daily-PP dataset.

多模态程序规划图文生成大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。