arXiv:2506.00625cs.CV2025-06TPAMI被引 2

通过融合头尾类别特征,提升长尾视觉识别准确率。

PI-H2T: Enhancing Long-Tailed Visual Recognition with Permutation-Invariant and Head-to-Tail Feature Fusion

  • 用置换不变特征融合增强表示空间聚类性。
  • 通过头尾信息传递改善尾部类别多样性,准确率显著提升。
  • 可即插即用,适合各类长尾识别模型优化。

长尾数据分布导致深度学习模型偏向头部类别而忽视尾部类别,主要原因是表示空间失真和分类器偏差,源于尾部类别语义信息不足。为此,我们提出置换不变且头尾特征融合(PI-H2T)方法:通过置换不变表示融合(PIF)优化表示空间,实现更紧密的特征聚集与自动类别间距;同时利用头尾融合(H2TF)将头部语义信息迁移至尾部类别,提升尾部多样性。理论分析与实验表明,PI-H2T同步优化了表示空间与决策边界。其即插即用设计可无缝集成至现有方法,为性能提升提供简单路径。在多个长尾基准测试上验证了其有效性。

原文摘要 · Abstract (English)

The imbalanced distribution of long-tailed data presents a significant challenge for deep learning models, causing them to prioritize head classes while neglecting tail classes. Two key factors contributing to low recognition accuracy are the deformed representation space and a biased classifier, stemming from insufficient semantic information in tail classes. To address these issues, we propose permutation-invariant and head-to-tail feature fusion (PI-H2T), a highly adaptable method. PI-H2T enhances the representation space through permutation-invariant representation fusion (PIF), yielding more clustered features and automatic class margins. Additionally, it adjusts the biased classifier by transferring semantic information from head to tail classes via head-to-tail fusion (H2TF), improving tail class diversity. Theoretical analysis and experiments show that PI-H2T optimizes both the representation space and decision boundaries. Its plug-and-play design ensures seamless integration into existing methods, providing a straightforward path to further performance improvements. Extensive experiments on long-tailed benchmarks confirm the effectiveness of PI-H2T.

长尾识别特征融合表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。