arXiv:2602.04152cs.ROcs.AI2026-02被引 2

多智能体协作生成大规模室内场景图,无需训练即可融合信息。

MA3DSG: Multi-Agent 3D Scene Graph Generation for Large-Scale Indoor Environments

  • 多个智能体并行采集局部场景图,通过无参对齐算法合成全局图。
  • 在真实规模环境中实现可扩展的场景图生成,支持多种配置组合。
  • 适合研究大规模三维场景理解与多智能体协同系统的设计者。

现有3D场景图生成方法普遍依赖单智能体假设和小规模环境,难以拓展至真实场景。本文提出首个面向大规模室内环境的多智能体3D场景图生成框架MA3DSG。我们设计了一种无需训练的图对齐算法,能够高效将各智能体产生的局部查询图合并为统一的全局场景图。基于大量分析与实证洞察,该方法使传统单智能体系统可在不引入任何可学习参数的前提下实现协同工作。为严格评估3DSGG性能,我们构建了MA3DSG-Bench——一个支持多样化智能体配置、领域规模及环境条件的基准测试平台,提供更具通用性与可扩展性的评估体系。本工作为可扩展的多智能体3DSGG研究奠定了坚实基础。

原文摘要 · Abstract (English)

Current 3D scene graph generation (3DSGG) approaches heavily rely on a single-agent assumption and small-scale environments, exhibiting limited scalability to real-world scenarios. In this work, we introduce Multi-Agent 3D Scene Graph Generation (MA3DSG) model, the first framework designed to tackle this scalability challenge using multiple agents. We develop a training-free graph alignment algorithm that efficiently merges partial query graphs from individual agents into a unified global scene graph. Leveraging extensive analysis and empirical insights, our approach enables conventional single-agent systems to operate collaboratively without requiring any learnable parameters. To rigorously evaluate 3DSGG performance, we propose MA3DSG-Bench-a benchmark that supports diverse agent configurations, domain sizes, and environmental conditions-providing a more general and extensible evaluation framework. This work lays a solid foundation for scalable, multi-agent 3DSGG research.

场景图生成多智能体3D理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。