arXiv:2606.24649cs.CV2026-06中稿 · ECCV

用多个智能体协作实现零样本3D理解,提升视角覆盖与认知一致性。

Agentic Collaborative Cognition for Zero-Shot 3D Understanding

论文配图:Agentic Collaborative Cognition for Zero-Shot 3D Understanding
图 1 · 摘自论文原文
  • 设计规划与感知双智能体,协同规划视角并构建统一3D认知地图。
  • 在ScanRefer、SQA3D等6个基准上达最新水平,最高提升14.6% BLEU-1。
  • 适合需要高精度3D理解的视觉语言任务,如交互式问答与场景解析。

近期研究将零样本3D理解重构为多模态大模型(MLLMs)下的视频关键帧理解任务。然而,现有方法受限于视频固有的观测视角有限性及对3D场景的隐式感知。本文提出一种协作多智能体框架:规划智能体负责高层视角规划并补充缺失视角,感知智能体则显式将3D场景整合为结构化认知地图。规划智能体分析该地图后确定相关视点并补充关键视角以确保全面观测;感知智能体从这些视点中分配一致实例标识符,记录物体级属性,并将碎片化信息融合进认知地图。同时提供反馈过滤错误候选对象,指导后续视角规划。通过闭环迭代,两智能体协同推断直至感知智能体确认任务完成。大量实验表明,本方法在6个基准上达到领先性能,包括ScanRefer上11.1% [email protected]提升,3D辅助对话任务中14.6% BLEU-1提升,以及SQA3D上2.1% EM提升。

原文摘要 · Abstract (English)

Recent advancements have explored agentic zero-shot 3D understanding by reformulating it as video keyframe understanding with Multimodal Large Language Models (MLLMs). However, existing methods face an intrinsic bottleneck due to the finite observation perspectives inherent in videos and the implicit perception of 3D scenes. In this paper, we propose a collaborative multi-agent framework that assigns a Planning Agent to handle high-level viewpoint planning and supplement novel perspectives, and a Perception Agent to explicitly summarize the 3D scene into a structured holistic cognitive map. Specifically, Planning Agent first analyzes this cognitive map to determine query-relevant viewpoints and supplements missing critical perspectives to ensure comprehensive observation. Subsequently, Perception Agent documents object-level attributes from these views by assigning consistent instance identifiers across viewpoints, thereby integrating fragmented observations into the holistic cognitive map. In parallel, it provides feedback to filter out mismatched candidate objects and guide subsequent viewpoint planning. Through this closed-loop iterative process, two agents collaboratively figure out candidates until Perception Agent determines that sufficient information has been captured to complete the task. Extensive experiments demonstrate that our method achieves state-of-the-art performance on 6 benchmarks, with improvements of 11.1\% [email protected] on ScanRefer, 14.6 BLEU-1 on 3D-assisted dialog, and 2.1 EM on SQA3D.

3D理解多智能体零样本认知地图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。