让AI同时理解图文指令,实时指导用户完成复杂操作。
Show and Guide: Instructional-Plan Grounded Vision and Language Model
- 融合文本与图像的多模态指令理解,支持图文交互。
- 可定位视频中对应步骤并生成下一步指引,准确率显著提升。
- 适合智能助手、机器人导航等需图文协同的任务场景。
指导用户完成复杂流程任务是典型的多模态任务,视觉化步骤信息对有效引导至关重要。然而,现有计划跟随语言模型(LMs)通常无法处理多模态输入输出。本文提出MM-PlanLLM,首个基于多模态指令计划的语言模型,通过文本计划和视觉信息协同,辅助用户执行任务。核心包含两项关键技术:对话式视频片段检索(Conversational Video Moment Retrieval),根据用户提问定位相关视频段;视觉感知步骤生成(Visually-Informed Step Generation),基于用户当前进度图像生成下一步指令。模型采用新颖的多任务多阶段训练策略,逐步学习多模态指令计划语义层,在多模态与纯文本对话任务上均表现优异。此外,验证了文本计划步骤与教学视频时刻之间在时序与结构上的跨模态对齐能力。
原文摘要 · Abstract (English)
Guiding users through complex procedural plans is an inherently multimodal task in which having visually illustrated plan steps is crucial to deliver an effective plan guidance. However, existing works on plan-following language models (LMs) often are not capable of multimodal input and output. In this work, we present MM-PlanLLM, the first multimodal LLM designed to assist users in executing instructional tasks by leveraging both textual plans and visual information. Specifically, we bring cross-modality through two key tasks: Conversational Video Moment Retrieval, where the model retrieves relevant step-video segments based on user queries, and Visually-Informed Step Generation, where the model generates the next step in a plan, conditioned on an image of the user's current progress. MM-PlanLLM is trained using a novel multitask-multistage approach, designed to gradually expose the model to multimodal instructional-plans semantic layers, achieving strong performance on both multimodal and textual dialogue in a plan-grounded setting. Furthermore, we show that the model delivers cross-modal temporal and plan-structure representations aligned between textual plan steps and instructional video moments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。