arXiv:2601.04404cs.CVcs.AI2026-01中稿 · NeurIPS

三模态协作框架提升大规模3D物体标注精度与效率

3D-Agent:Tri-Modal Multi-Agent Collaboration for Scalable 3D Object Annotation

  • 融合2D多视角图像、文本描述和3D点云的三模态协同标注
  • 在多个数据集上达到88.7的CLIPScore,每小时处理1.2万个物体
  • 适合自动驾驶、机器人等需要高精度3D标注的场景

自动驾驶、机器人和增强现实应用推动了3D物体标注的发展,但其面临空间复杂性、遮挡和视点不一致等挑战。现有基于单模型的方法难以有效应对这些问题。本文提出Tri MARF,一种集成三模态输入(2D多视角图像、文本描述、3D点云)的多智能体协作框架,用于提升大规模3D标注效果。该框架包含三个专用智能体:视觉语言模型智能体生成多视角描述,信息聚合智能体选择最优描述,门控智能体将文本语义与3D几何对齐以实现精细化描述。在Objaverse、LVIS、Objaverse XL和ABO数据集上的大量实验表明,Tri MARF显著优于现有方法,达到88.7的CLIPScore,检索准确率分别在ViLT R@5上达45.2和43.8,并在单张NVIDIA A100 GPU上实现最高每小时12000个物体的吞吐量。

原文摘要 · Abstract (English)

Driven by applications in autonomous driving robotics and augmented reality 3D object annotation presents challenges beyond 2D annotation including spatial complexity occlusion and viewpoint inconsistency Existing approaches based on single models often struggle to address these issues effectively We propose Tri MARF a novel framework that integrates tri modal inputs including 2D multi view images textual descriptions and 3D point clouds within a multi agent collaborative architecture to enhance large scale 3D annotation Tri MARF consists of three specialized agents a vision language model agent for generating multi view descriptions an information aggregation agent for selecting optimal descriptions and a gating agent that aligns textual semantics with 3D geometry for refined captioning Extensive experiments on Objaverse LVIS Objaverse XL and ABO demonstrate that Tri MARF substantially outperforms existing methods achieving a CLIPScore of 88 point 7 compared to prior state of the art methods retrieval accuracy of 45 point 2 and 43 point 8 on ViLT R at 5 and a throughput of up to 12000 objects per hour on a single NVIDIA A100 GPU

3D标注多模态智能体协作自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。