arXiv:2604.09167cs.CVcs.MA2026-04

用多个智能体协作,让模型在3D场景中零样本理解并推理物体位置关系。

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding

  • 设计三个专家智能体:规划、定位、编程,动态协作完成3D推理
  • 无需训练,在复杂场景下零样本达到领先性能
  • 适合需要灵活推理3D环境的机器人、AR/VR应用

视觉语言模型(VLMs)在多模态理解与推理上表现优异,但三维场景中的具身推理仍待深入。有效3D推理依赖精准定位:回答开放式问题前,需先识别场景中相关物体与区域,并分析其空间和几何关系。现有方法虽有潜力,但常依赖领域内微调或人工设计的推理流程,限制了灵活性与零样本泛化能力。本文提出MAG-3D,一个基于现成VLM的无训练多智能体框架,用于3D具身推理。通过动态协调专家智能体,解决关键挑战:规划智能体分解任务并统筹推理流程;定位智能体从海量3D观测中进行自由形式定位与相关帧检索;编码智能体通过可执行程序实现灵活几何推理与显式验证。该协同设计使模型在多样场景中实现灵活、无训练的3D具身推理,并在挑战性基准上达到当前最优表现。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have achieved strong performance in multimodal understanding and reasoning, yet grounded reasoning in 3D scenes remains underexplored. Effective 3D reasoning hinges on accurate grounding: to answer open-ended queries, a model must first identify query-relevant objects and regions in a complex scene, and then reason about their spatial and geometric relationships. Recent approaches have demonstrated strong potential for grounded 3D reasoning. However, they often rely on in-domain tuning or hand-crafted reasoning pipelines, which limit their flexibility and zero-shot generalization to novel environments. In this work, we present MAG-3D, a training-free multi-agent framework for grounded 3D reasoning with off-the-shelf VLMs. Instead of relying on task-specific training or fixed reasoning procedures, MAG-3D dynamically coordinates expert agents to address the key challenges of 3D reasoning. Specifically, we propose a planning agent that decomposes the task and orchestrates the overall reasoning process, a grounding agent that performs free-form 3D grounding and relevant frame retrieval from extensive 3D scene observations, and a coding agent that conducts flexible geometric reasoning and explicit verification through executable programs. This multi-agent collaborative design enables flexible training-free 3D grounded reasoning across diverse scenes and achieves state-of-the-art performance on challenging benchmarks.

3D理解多智能体具身推理零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。