提升大模型在复杂3D场景中准确定位目标物的能力
3D Spatial Understanding in MLLMs: Disambiguation and Evaluation
- 通过简单有效方法增强模型对目标物的空间定位能力
- 在3D视觉定位任务上表现优于现有方法,实现更高精度
- 适合需要精准空间指令的机器人协作场景
多模态大语言模型(MLLMs)在图像描述和问答等任务上已取得显著进展。然而,尽管能生成逼真的描述,它们在复杂3D环境中对物体进行精确定位和消歧方面仍存在困难。这一能力对与协作机器人系统集成至关重要。当目标物体被相似物体包围时,机器人需提供清晰、具备空间感知的指令以有效引导人类。我们称此挑战为上下文相关的物体定位与消歧,其要求严于传统的3D密集描述任务,尤其强调目标的唯一性。为此,我们提出简单而有效的技术,增强模型定位和消歧目标物体的能力。该方法不仅在评估句法相似度的传统指标上达到领先水平,还通过3D视觉定位模型验证了其在3D空间理解上的显著提升。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have made significant progress in tasks such as image captioning and question answering. However, while these models can generate realistic captions, they often struggle with providing precise instructions, particularly when it comes to localizing and disambiguating objects in complex 3D environments. This capability is critical as MLLMs become more integrated with collaborative robotic systems. In scenarios where a target object is surrounded by similar objects (distractors), robots must deliver clear, spatially-aware instructions to guide humans effectively. We refer to this challenge as contextual object localization and disambiguation, which imposes stricter constraints than conventional 3D dense captioning, especially regarding ensuring target exclusivity. In response, we propose simple yet effective techniques to enhance the model's ability to localize and disambiguate target objects. Our approach not only achieves state-of-the-art performance on conventional metrics that evaluate sentence similarity, but also demonstrates improved 3D spatial understanding through 3D visual grounding model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。