用原型学习提升细粒度3D形状分类的准确率与可解释性
Proto-FG3D: Prototype-based Interpretable Fine-Grained 3D Shape Classification
- 基于原型的非参数学习替代传统分类器,实现多视角联合表征
- 通过在线聚类优化原型,提升细粒度区分能力与类别平衡性
- 支持透明推理,适合需要可解释性的3D形状识别场景
基于深度学习的多视角粗粒度3D形状分类在过去十年取得了显著进展,得益于CNN和ViT等骨干网络强大的特征学习能力。然而,对于需要精细理解形状的细粒度3D分类任务,仍因多视角特征聚合中判别信息不足、类间微小差异难以捕捉、类别不平衡及参数化模型固有的可解释性局限而研究较少。为此,本文提出首个原型驱动的框架Proto-FG3D,实现从参数化softmax到非参数原型学习的范式转变。首先,通过原型关联建立多视角与多类别联合表征;其次,利用在线聚类优化原型,增强多视角特征分配鲁棒性与子类间平衡性;最后,基于原型-视角相关性分析构建监督学习机制,提升细粒度判别能力,并通过透明的案例推理实现即兴可解释性。在FG3D和ModelNet40数据集上的实验表明,Proto-FG3D在准确率、透明预测与即兴可解释性方面均超越现有方法,挑战了传统的细粒度3D识别范式。
原文摘要 · Abstract (English)
Deep learning-based multi-view coarse-grained 3D shape classification has achieved remarkable success over the past decade, leveraging the powerful feature learning capabilities of CNN-based and ViT-based backbones. However, as a challenging research area critical for detailed shape understanding, fine-grained 3D classification remains understudied due to the limited discriminative information captured during multi-view feature aggregation, particularly for subtle inter-class variations, class imbalance, and inherent interpretability limitations of parametric model. To address these problems, we propose the first prototype-based framework named Proto-FG3D for fine-grained 3D shape classification, achieving a paradigm shift from parametric softmax to non-parametric prototype learning. Firstly, Proto-FG3D establishes joint multi-view and multi-category representation learning via Prototype Association. Secondly, prototypes are refined via Online Clustering, improving both the robustness of multi-view feature allocation and inter-subclass balance. Finally, prototype-guided supervised learning is established to enhance fine-grained discrimination via prototype-view correlation analysis and enables ad-hoc interpretability through transparent case-based reasoning. Experiments on FG3D and ModelNet40 show Proto-FG3D surpasses state-of-the-art methods in accuracy, transparent predictions, and ad-hoc interpretability with visualizations, challenging conventional fine-grained 3D recognition approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。