用长上下文建模提升机器人对话中的动作确认与规划能力
Robot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLM
- 引入双侧上下文依赖的长序列Q-former,捕捉视频全程动作关联
- 在YouCook2数据集上,动作确认准确率显著提升,带动整体规划性能改善
- 结合文本条件输入优化大模型信息表达,适合复杂人机协作场景
人类与机器人协同完成共同目标,需要机器人理解人类行为及与环境的互动。本文聚焦基于对话的人机交互,利用多模态场景理解实现机器人动作确认与步骤生成。现有先进方法使用多模态变换器,从单个片段中生成与动作确认对齐的动作步骤,但这些方法主要依赖片段级处理,未能充分利用完整视频中的长程上下文信息。为此,本文提出一种融合左右上下文依赖的长上下文Q-former,增强对长时间任务中动作间依赖关系的建模能力。同时,提出文本条件化方法,将文本嵌入直接输入LLM解码器,缓解Q-former对文本信息的过度抽象问题。在YouCook2语料库上的实验表明,动作确认生成准确率是影响动作规划性能的关键因素;进一步验证了长上下文Q-former结合VideoLLaMA3可有效提升确认与规划效果。
原文摘要 · Abstract (English)
Human-robot collaboration towards a shared goal requires robots to understand human action and interaction with the surrounding environment. This paper focuses on human-robot interaction (HRI) based on human-robot dialogue that relies on the robot action confirmation and action step generation using multimodal scene understanding. The state-of-the-art approach uses multimodal transformers to generate robot action steps aligned with robot action confirmation from a single clip showing a task composed of multiple micro steps. Although actions towards a long-horizon task depend on each other throughout an entire video, the current approaches mainly focus on clip-level processing and do not leverage long-context information. This paper proposes a long-context Q-former incorporating left and right context dependency in full videos. Furthermore, this paper proposes a text-conditioning approach to feed text embeddings directly into the LLM decoder to mitigate the high abstraction of the information in text by Q-former. Experiments with the YouCook2 corpus show that the accuracy of confirmation generation is a major factor in the performance of action planning. Furthermore, we demonstrate that the long-context Q-former improves the confirmation and action planning by integrating VideoLLaMA3.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。