arXiv:2503.10529cs.CVcs.AI2025-03被引 9

用自增强方法生成高质量3D图文数据,提升大模型理解能力

PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Models

  • 融合2D与3D大模型优势,循环生成带空间语义的3D指令数据
  • 新模型在零样本3D描述和生成分类上分别提升8.33%和16.25%
  • 构建覆盖六大维度的新基准PiSA-Bench,更全面评估3D理解能力

3D多模态大语言模型虽有进展,但受限于3D数据量少、质量差。现有方法从2D模型迁移知识,仍存在模态和领域差距。为此,我们提出PiSA-Engine(点云自增强引擎),通过整合现成3D与2D大模型的全局洞察,实现高质量3D点云-文本指令数据的持续生成。以PointLLM为基础,采用协同进化训练框架,构建增强版3D MLLM——PointLLM-PiSA。同时,发现以往3D基准存在语言描述粗略、类别覆盖不足等问题,导致评估失真,因此引入新基准PiSA-Bench,涵盖六项关键维度并提供详细多样标签。实验表明,PointLLM-PiSA在零样本3D物体描述和生成分类任务上分别取得46.45%(+8.33%)和63.75%(+16.25%)的显著提升。代码、数据集与基准将开源。

原文摘要 · Abstract (English)

3D Multimodal Large Language Models (MLLMs) have recently made substantial advancements. However, their potential remains untapped, primarily due to the limited quantity and suboptimal quality of 3D datasets. Current approaches attempt to transfer knowledge from 2D MLLMs to expand 3D instruction data, but still face modality and domain gaps. To this end, we introduce PiSA-Engine (Point-Self-Augmented-Engine), a new framework for generating instruction point-language datasets enriched with 3D spatial semantics. We observe that existing 3D MLLMs offer a comprehensive understanding of point clouds for annotation, while 2D MLLMs excel at cross-validation by providing complementary information. By integrating holistic 2D and 3D insights from off-the-shelf MLLMs, PiSA-Engine enables a continuous cycle of high-quality data generation. We select PointLLM as the baseline and adopt this co-evolution training framework to develop an enhanced 3D MLLM, termed PointLLM-PiSA. Additionally, we identify limitations in previous 3D benchmarks, which often feature coarse language captions and insufficient category diversity, resulting in inaccurate evaluations. To address this gap, we further introduce PiSA-Bench, a comprehensive 3D benchmark covering six key aspects with detailed and diverse labels. Experimental results demonstrate PointLLM-PiSA's state-of-the-art performance in zero-shot 3D object captioning and generative classification on our PiSA-Bench, achieving significant improvements of 46.45% (+8.33%) and 63.75% (+16.25%), respectively. We will release the code, datasets, and benchmark.

3D理解大模型数据增强多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。