arXiv:2511.08971cs.HCcs.CV2025-11被引 3

零样本解决第一人称视角意图模糊问题,提升交互准确性。

Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation

  • 分模块设计三类澄清器,分别处理语言、视觉和跨模态歧义。
  • 小模型性能提升约30%,大模型也获稳定增益,视觉引导准确率提高20%以上。
  • 无需训练即可部署,适合增强智能眼镜等可穿戴设备的交互体验。

第一人称视角人工智能代理的性能受限于多模态意图模糊性。这种挑战源于语言描述不明确、视觉数据不完整以及指示性手势,常导致任务失败。现有单一架构的视觉-语言模型难以解决此类多模态模糊输入,往往无声失败或生成幻觉响应。为此,我们提出即插即用澄清器(Plug-and-Play Clarifier),一种零样本、模块化框架,将问题分解为可解的子任务。该框架包含三个协同模块:(1) 文本澄清器,通过对话式推理互动消除语言意图歧义;(2) 视觉澄清器,实时提供引导反馈,指导用户调整姿态以提升捕捉质量;(3) 跨模态澄清器,具备定位机制,能鲁棒地解析3D指向手势并识别用户所指的具体物体。大量实验表明,该框架使小型语言模型(4–8B)的意图澄清性能提升约30%,使其在性能上接近更大规模模型。对大模型亦观察到一致增益。此外,视觉澄清器使纠正引导准确率提升超20%,跨模态澄清器使指代定位语义回答准确率提升5%。整体而言,该方法提供了一种可即插即用的框架,有效缓解多模态模糊性,显著改善第一人称交互体验。

原文摘要 · Abstract (English)

The performance of egocentric AI agents is fundamentally limited by multimodal intent ambiguity. This challenge arises from a combination of underspecified language, imperfect visual data, and deictic gestures, which frequently leads to task failure. Existing monolithic Vision-Language Models (VLMs) struggle to resolve these multimodal ambiguous inputs, often failing silently or hallucinating responses. To address these ambiguities, we introduce the Plug-and-Play Clarifier, a zero-shot and modular framework that decomposes the problem into discrete, solvable sub-tasks. Specifically, our framework consists of three synergistic modules: (1) a text clarifier that uses dialogue-driven reasoning to interactively disambiguate linguistic intent, (2) a vision clarifier that delivers real-time guidance feedback, instructing users to adjust their positioning for improved capture quality, and (3) a cross-modal clarifier with grounding mechanism that robustly interprets 3D pointing gestures and identifies the specific objects users are pointing to. Extensive experiments demonstrate that our framework improves the intent clarification performance of small language models (4--8B) by approximately 30%, making them competitive with significantly larger counterparts. We also observe consistent gains when applying our framework to these larger models. Furthermore, our vision clarifier increases corrective guidance accuracy by over 20%, and our cross-modal clarifier improves semantic answer accuracy for referential grounding by 5%. Overall, our method provides a plug-and-play framework that effectively resolves multimodal ambiguity and significantly enhances user experience in egocentric interaction.

第一人称多模态意图识别交互系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。