用对抗与多样指令数据训练3D大模型,提升理解与泛化能力。
Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning
- 构建对抗与多样指令数据集,增强模型判别力
- 在多3D任务上实现7.8%和6.9%的性能提升
- 无需微调即可适配多种3D多模态任务
3D大语言模型(3DLLMs)在构建三维现实世界通用智能体方面展现出巨大潜力,但受限于高质量鲁棒指令跟随数据的缺乏,其判别能力和泛化性能仍受限。本文提出Robin3D,一种基于全新数据引擎Robust Instruction Generation(RIG)生成的大规模指令跟随数据训练的3DLLM。RIG生成两类关键数据:1)对抗性指令数据,包含正负样本混合,提升模型判别理解能力;2)多样化指令数据,涵盖多种指令风格,增强模型泛化能力。最终构建100万条指令数据,包括34.4万对抗样本、50.8万多样化样本及16.5万基准训练集样本。为更好处理复杂指令,Robin3D首先引入关系增强投影器以强化空间理解,再通过ID-特征绑定提升物体指代与定位能力。Robin3D在五个主流3D多模态学习基准上持续超越先前方法,无需任务特定微调。显著成果包括在定位任务(Multi3DRefer)中提升7.8%,在描述生成任务(Scan2Cap)中提升6.9%。
原文摘要 · Abstract (English)
Recent advancements in 3D Large Language Models (3DLLMs) have highlighted their potential in building general-purpose agents in the 3D real world, yet challenges remain due to the lack of high-quality robust instruction-following data, leading to limited discriminative power and generalization of 3DLLMs. In this paper, we introduce Robin3D, a powerful 3DLLM trained on large-scale instruction-following data generated by our novel data engine, Robust Instruction Generation (RIG) engine. RIG generates two key instruction data: 1) the Adversarial Instruction-following data, which features mixed negative and positive samples to enhance the model's discriminative understanding. 2) the Diverse Instruction-following data, which contains various instruction styles to enhance model's generalization. As a result, we construct 1 million instruction-following data, consisting of 344K Adversarial samples, 508K Diverse samples, and 165K benchmark training set samples. To better handle these complex instructions, Robin3D first incorporates Relation-Augmented Projector to enhance spatial understanding, and then strengthens the object referring and grounding ability through ID-Feature Bonding. Robin3D consistently outperforms previous methods across five widely-used 3D multimodal learning benchmarks, without the need for task-specific fine-tuning. Notably, we achieve a 7.8\% improvement in the grounding task (Multi3DRefer) and a 6.9\% improvement in the captioning task (Scan2Cap).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。