用眼神和语音增强任务演示,让AI更好理解人类行为意图。
Grounding Task Assistance with Multimodal Cues from a Single Demonstration

- 融合眼动与语音信号,从单次演示中提取细粒度意图
- 结合多模态线索,问答准确率显著高于仅靠图像帧
- 适合需要个性化交互的智能助手场景
人类演示常作为他人学习任务的关键参考。然而,主流的RGB视频难以捕捉行为中蕴含的意图、安全关键环境因素及细微偏好等细粒度上下文线索。这一感知差距严重限制了视觉语言模型(VLMs)对行为原因的理解与个性化适应能力。为此,我们提出MICA(Multimodal Interactive Contextualized Assistance)框架,通过整合眼动与语音线索,提升对话式任务助手的表现。MICA将演示分解为有意义的子任务,提取关键帧与描述,实现更丰富的上下文语义定位,支持视觉问答。在实时聊天辅助任务复现中,基于多模态线索的问答表现显著优于仅依赖图像帧的检索方法。值得注意的是,仅使用眼动线索即可达到语音性能的93%,两者结合时准确率最高。任务类型决定了隐式(眼动)与显式(语音)线索的有效性,凸显了自适应多模态模型的重要性。结果表明,基于帧的上下文存在局限,多模态信号对真实世界AI任务辅助具有关键价值。
原文摘要 · Abstract (English)
A person's demonstration often serves as a key reference for others learning the same task. However, RGB video, the dominant medium for representing these demonstrations, often fails to capture fine-grained contextual cues such as intent, safety-critical environmental factors, and subtle preferences embedded in human behavior. This sensory gap fundamentally limits the ability of Vision Language Models (VLMs) to reason about why actions occur and how they should adapt to individual users. To address this, we introduce MICA (Multimodal Interactive Contextualized Assistance), a framework that improves conversational agents for task assistance by integrating eye gaze and speech cues. MICA segments demonstrations into meaningful sub-tasks and extracts keyframes and captions that capture fine-grained intent and user-specific cues, enabling richer contextual grounding for visual question answering. Evaluations on questions derived from real-time chat-assisted task replication show that multimodal cues significantly improve response quality over frame-based retrieval. Notably, gaze cues alone achieves 93% of speech performance, and their combination yields the highest accuracy. Task type determines the effectiveness of implicit (gaze) vs. explicit (speech) cues, underscoring the need for adaptable multimodal models. These results highlight the limitations of frame-based context and demonstrate the value of multimodal signals for real-world AI task assistance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。