评测并提升视觉智能眼镜对指向动作的理解能力
Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision

- 构建涵盖1.1万+样本的指向理解基准测试集
- 发现主流模型常误判指向目标,存在空间错觉现象
- 通过合成数据微调,显著提升真实场景泛化能力
以智能眼镜为代表的自指视觉人工智能系统依赖指向手势来解析自然语言指令中的指代歧义。然而,尽管多模态大模型(MLLMs)持续发展,当前系统仍难以准确捕捉指向的空间语义,往往依赖视觉邻近或物体显著性等表面关联,这种现象被称为“指代幻觉”。为此,本文提出EgoPoint-Bench,一个涵盖超过11,000个高保真模拟与真实世界样本的问答基准,覆盖五个评估维度和三个指代复杂度层级。大量实验表明,尽管最先进的专有与开源模型在自指指向任务中表现不佳,但经过我们合成数据微调的模型展现出显著性能提升,并具备稳健的模拟到真实场景泛化能力。本工作强调空间感知监督的重要性,为实现精准自指智能助手提供可扩展路径。
原文摘要 · Abstract (English)
Egocentric AI agents, such as smart glasses, rely on pointing gestures to resolve referential ambiguities in natural language commands. However, despite advancements in Multimodal Large Language Models (MLLMs), current systems often fail to precisely ground the spatial semantics of pointing. Instead, they rely on spurious correlations with visual proximity or object saliency, a phenomenon we term "Referential Hallucination." To address this gap, we introduce EgoPoint-Bench, a comprehensive question-answering benchmark designed to evaluate and enhance multimodal pointing reasoning in egocentric views. Comprising over 11k high-fidelity simulated and real-world samples, the benchmark spans five evaluation dimensions and three levels of referential complexity. Extensive experiments demonstrate that while state-of-the-art proprietary and open-source models struggle with egocentric pointing, models fine-tuned on our synthetic data achieve significant performance gains and robust sim-to-real generalization. This work highlights the importance of spatially aware supervision and offers a scalable path toward precise egocentric AI assistants. Project page: https://guyyyug.github.io/EgoPoint-Bench/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。