通过视觉图增强,提升大模型在第一视角场景中的空间理解能力。
Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation

- 构建第一视角元素图作为中间表示,融合视觉模型的空间特征。
- 在室内和室外场景中分别提升8.14%和8.72%的问答准确率。
- 尤其在购物场景表现突出,适合需要精准空间推理的应用。
第一视角视觉问答(Ego VQA)是使多模态大语言模型(MLLMs)与现实世界交互的重要任务。然而,现有MLLMs在复杂第一视角场景中因空间感知能力有限,难以有效进行空间推理。为此,我们提出第一视角场景增强框架(ESA),通过所提出的“第一视角元素图”主动增强模型从第一人称视角出发的空间感知能力。核心思路是利用该图作为中介表示,借助视觉基础模型来增强MLLMs对第一视角场景的空间理解。具体包括:1)构建第一视角元素图,整合由视觉基础模型提取的第一视角空间特征;2)通过该图增强MLLMs在第一视角场景中的空间感知能力。在EgoTextVQA基准测试中,我们的方法取得显著性能提升:室内场景提升8.14%,室外场景提升8.72%。尤其在室内场景的购物子集上表现最为突出。项目代码已开源。
原文摘要 · Abstract (English)
Egocentric Visual Question Answering (VQA) has attracted widespread attention as an important task for enabling Multimodal Large Language Models (MLLMs) to interact with the real world. However, existing MLLMs struggle to perform effective spatial reasoning in complex egocentric scenes due to their limited spatial perception capabilities. To this end, we introduce Ego Scene Augmentation (ESA), an egocentric spatial perception framework, which actively enhances the spatial perception capabilities from the egocentric perspective, powered by the proposed Ego-element Graph. Our core insight is leveraging the Ego-element Graph as an intermediary representation to augment the egocentric spatial perception of MLLMs via visual foundational models. Specifically, we 1) construct the Ego-element Graph, which encapsulates and integrates egocentric spatial features enabled by visual foundational models; 2) enhance the spatial perception capabilities of MLLMs via the Ego-element Graph for ego-perspective scenes. Our proposed ESA framework presents significant performance improvement on the EgoTextVQA benchmark. We achieve an 8.14% gain on the indoor setting and an 8.72% gain on the outdoor setting. Furthermore, our ESA shows the most impressive performance improvement in the shopping subset of the indoor setting. The project code is publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。