用文本生成3D数据扩增,提升零样本3D识别效果
Text-guided Synthetic Geometric Augmentation for Zero-shot 3D Understanding
- 通过文本引导生成3D数据并过滤语义不一致样本
- 在三个数据集上零样本准确率提升3.0%至8.7%
- 适合缺乏真实3D数据的零样本3D视觉研究者
零样本3D分类需要大量训练数据,但收集3D数据和描述成本高,远超2D视觉。生成模型在合成数据方面取得突破,为解决此问题带来可能。本文提出文本引导几何增强(TeGA)方法,专为语言-图像-3D预训练设计,利用生成式文本到3D模型扩展有限3D数据集。具体地,自动生成文本引导的合成3D数据,并引入一致性过滤策略,剔除语义与几何形状不符的噪声样本。实验中将原始数据集规模翻倍,相较基线模型在Objaverse-LVIS上提升3.0%,ScanObjectNN上提升4.6%,ModelNet40上提升8.7%,证明TeGA能有效缓解3D数据短缺,即使在真实数据有限下仍可实现鲁棒的零样本3D分类,推动零样本3D视觉应用落地。
原文摘要 · Abstract (English)
Zero-shot recognition models require extensive training data for generalization. However, in zero-shot 3D classification, collecting 3D data and captions is costly and laborintensive, posing a significant barrier compared to 2D vision. Recent advances in generative models have achieved unprecedented realism in synthetic data production, and recent research shows the potential for using generated data as training data. Here, naturally raising the question: Can synthetic 3D data generated by generative models be used as expanding limited 3D datasets? In response, we present a synthetic 3D dataset expansion method, Textguided Geometric Augmentation (TeGA). TeGA is tailored for language-image-3D pretraining, which achieves SoTA in zero-shot 3D classification, and uses a generative textto-3D model to enhance and extend limited 3D datasets. Specifically, we automatically generate text-guided synthetic 3D data and introduce a consistency filtering strategy to discard noisy samples where semantics and geometric shapes do not match with text. In the experiment to double the original dataset size using TeGA, our approach demonstrates improvements over the baselines, achieving zeroshot performance gains of 3.0% on Objaverse-LVIS, 4.6% on ScanObjectNN, and 8.7% on ModelNet40. These results demonstrate that TeGA effectively bridges the 3D data gap, enabling robust zero-shot 3D classification even with limited real training data and paving the way for zero-shot 3D vision application.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。