用多模态大模型实现实时任务辅助,能看懂用户操作并主动提供帮助。
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models
- 结合视觉流与文本数据训练多模态模型,实现任务理解。
- 从视频自动构建任务图谱,在训练和推理中提升准确率。
- 在任务识别、动作预测等四项任务上表现领先,适合智能助手研发者。
生成式模型能力的提升使得构建超越语言的多模态虚拟助手成为可能。通过观察人类执行多步骤任务,可构建具备情境感知能力的助手,根据当前任务状态提供针对性协助。本文提出一种基于多模态大语言模型的上下文感知任务助手InsTALL,利用在线视觉流(如屏幕共享或视频录制)实时响应用户关于当前任务的提问。为实现有效协助,InsTALL 1)在任务视频与配对文本数据上训练多模态模型;2)自动从视频数据中提取任务图谱,并在训练和推理阶段加以利用。实验表明,InsTALL在提出的多模态活动理解子任务——任务识别(TR)、动作识别(AR)、下一步动作预测(AP)和计划预测(PP)上达到当前最优性能,并在两个新提出的自动错误识别子任务上优于现有基线。
原文摘要 · Abstract (English)
The improved competence of generative models can help building multi-modal virtual assistants that leverage modalities beyond language. By observing humans performing multi-step tasks, one can build assistants that have situational awareness of actions and tasks being performed, enabling them to cater assistance based on this understanding. In this paper, we develop a Context-aware Instructional Task Assistant with Multi-modal Large Language Models (InsTALL) that leverages an online visual stream (e.g. a user's screen share or video recording) and responds in real-time to user queries related to the task at hand. To enable useful assistance, InsTALL 1) trains a multi-modal model on task videos and paired textual data, and 2) automatically extracts task graph from video data and leverages it at training and inference time. We show InsTALL achieves state-of-the-art performance across proposed sub-tasks considered for multimodal activity understanding -- task recognition (TR), action recognition (AR), next action prediction (AP), and plan prediction (PP) -- and outperforms existing baselines on two novel sub-tasks related to automatic error identification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。