让3D视觉语言模型自动聚焦关键区域,提升大场景理解能力
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences
- 用大模型视觉偏好动态识别任务相关区域
- 通过可插拔的场景放大模块捕捉细节,性能显著提升
- 适合需要精准大场景理解的机器人导航与问答系统
3D视觉语言模型(3D-VLMs)在具身智能任务如视觉导航和具身问答中至关重要。由于大场景中视觉特征密集,准确定位任务相关区域极具挑战。现有方法对所有物体进行分割并提取特征,但这些无任务特性的特征包含大量冗余信息且缺乏任务相关区域的细节。为此,我们提出LSceneLLM,一种自适应框架:利用大语言模型(LLM)对不同任务的视觉偏好,自动识别任务相关区域,并通过即插即用的场景放大模块捕捉聚焦区域的细粒度特征。具体地,密集标记选择器分析LLM的注意力图,识别指令输入对应的视觉偏好,进而放大关注区域的细节。自适应自注意力模块融合粗粒度与精选细粒度视觉信息。为全面评估3D-VLM的大场景理解能力,我们还引入跨房间理解基准XR-Scene,包含XR-QA、XR-EmbodiedPlanning和XR-SceneCaption等任务。实验表明,该方法在大场景理解及现有基准上均优于现有方法,将场景放大模块嵌入现有3D-VLM也带来显著提升。
原文摘要 · Abstract (English)
Research on 3D Vision-Language Models (3D-VLMs) is gaining increasing attention, which is crucial for developing embodied AI within 3D scenes, such as visual navigation and embodied question answering. Due to the high density of visual features, especially in large 3D scenes, accurately locating task-relevant visual information is challenging. Existing works attempt to segment all objects and consider their features as scene representations. However, these task-agnostic object features include much redundant information and missing details for the task-relevant area. To tackle these problems, we propose LSceneLLM, an adaptive framework that automatically identifies task-relevant areas by leveraging LLM's visual preference for different tasks, followed by a plug-and-play scene magnifier module to capture fine-grained details in focused areas. Specifically, a dense token selector examines the attention map of LLM to identify visual preferences for the instruction input. It then magnifies fine-grained details of the focusing area. An adaptive self-attention module is leveraged to fuse the coarse-grained and selected fine-grained visual information. To comprehensively evaluate the large scene understanding ability of 3D-VLMs, we further introduce a cross-room understanding benchmark, XR-Scene, which contains a series of large scene understanding tasks including XR-QA, XR-EmbodiedPlanning, and XR-SceneCaption. Experiments show that our method surpasses existing methods on both large scene understanding and existing scene understanding benchmarks. Plunging our scene magnifier module into the existing 3D-VLMs also brings significant improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。