arXiv:2501.17547cs.CV2025-01被引 1

用3D生成模型实现无需训练的开放世界分类,支持新类别和任意姿态识别。

Towards Training-Free Open-World Classification with 3D Generative Models

  • 基于3D生成模型构建旋转不变特征提取器,无需训练
  • 在ModelNet10和McGill上分别提升32.0%和8.7%准确率
  • 适合动态场景中快速部署的新类别识别任务

3D开放世界分类在动态非结构化现实场景中具有挑战性且至关重要,需同时实现开放类别与开放姿态识别。现有方法多依赖复杂的2D预训练模型以获取丰富稳定的表征,但其性能受限于3D物体向2D空间投影的难题。本文首次探索使用3D生成模型解决该问题,利用其蕴含的先验知识,设计旋转不变特征提取器。该创新组合使系统具备免训练、开放类别和姿态不变特性,适用于3D开放世界分类。在基准数据集上的大量实验表明,该方法在ModelNet10和McGill上分别取得32.0%和8.7%的准确率提升,达到当前最优性能。

原文摘要 · Abstract (English)

3D open-world classification is a challenging yet essential task in dynamic and unstructured real-world scenarios, requiring both open-category and open-pose recognition. To address these challenges, recent wisdom often takes sophisticated 2D pre-trained models to provide enriched and stable representations. However, these methods largely rely on how 3D objects can be projected into 2D space, which is unfortunately not well solved, and thus significantly limits their performance. Unlike these present efforts, in this paper we make a pioneering exploration of 3D generative models for 3D open-world classification. Drawing on abundant prior knowledge from 3D generative models, we additionally craft a rotation-invariant feature extractor. This innovative synergy endows our pipeline with the advantages of being training-free, open-category, and pose-invariant, thus well suited to 3D open-world classification. Extensive experiments on benchmark datasets demonstrate the potential of generative models in 3D open-world classification, achieving state-of-the-art performance on ModelNet10 and McGill with 32.0% and 8.7% overall accuracy improvement, respectively.

3D分类生成模型开放世界免训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。