用视频语言模型实现开放词汇意图识别,准确率达80%
IntentVLM: Open-Vocabulary Intention Recognition through Forward-Inverse Modeling with Video-Language Models

- 分两阶段生成意图候选并筛选,减少推理幻觉
- 在两个数据集上达80%准确率,比基线高30%
- 适合需要理解复杂人类意图的机器人系统
提升人机交互效果需让机器人准确理解人类目标。在多模态场景中,机器人须融合文本、视觉等异构信号形成连贯意图判断。本文提出IntentVLM,一种基于认知科学前向-逆向建模思想的两阶段视频语言框架,将意图理解分解为候选生成与结构化筛选,有效降低隐空间推理中的幻觉。在IntentQA和Inst-IT Bench数据集上评估,其准确率最高达80%,显著超越基线30%,接近人类表现。结果表明,该结构化推理方法在不产生灾难性遗忘的前提下,提升了开放词汇意图理解能力,为人机协同机器人提供可靠基础。
原文摘要 · Abstract (English)
Improving the effectiveness of human-robot interaction requires social robots to accurately infer human goals through robust intention understanding. This challenge is particularly critical in multimodal settings, where agents must integrate heterogeneous signals including text, visual cues to form a coherent interpretation of user intent. This paper presents IntentVLM, a novel two-stage video-language framework designed for open-vocabulary human intention recognition. The approach is inspired by forward-inverse modeling in cognitive science by decomposing intention understanding into goal candidate generation followed by structured inference through selection, effectively reducing hallucinations in latent reasoning. Evaluated on the IntentQA and Inst-IT Bench datasets, IntentVLM achieves state-of-the-art results with up to 80% accuracy, notably surpassing the baseline performance by 30% and matches human performance. Our findings demonstrate that this structured reasoning approach enhances open-vocabulary intention understanding without catastrophic forgetting, offering a robust foundation for human-centered robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。