arXiv:2512.06345cs.CV2025-12AAAI

让神经网络具备可解释的聚类注意力机制,提升视觉理解效果

CLUENet: Cluster Attention Makes Neural Networks Have Eyes

  • 用温度缩放余弦注意力与门控残差连接增强局部建模
  • 在块间实现硬分配与共享特征调度,提升效率与性能
  • 改进聚类池化策略,兼顾准确率、速度与模型透明性

尽管卷积与注意力模型在视觉任务中表现优异,但其固定的感受野和复杂架构限制了对不规则空间模式的建模能力,并影响可解释性,不利于需要高透明度的任务。聚类范式虽具良好可解释性和灵活语义建模能力,却存在精度不足、效率低及训练中梯度消失等问题。为此,我们提出CLUster attEntion Network(CLUENet),一种面向视觉语义理解的透明深度架构。主要创新包括:(i) 全局软聚合与硬分配结合的温度缩放余弦注意力,辅以门控残差连接,增强局部建模;(ii) 块间硬分配与共享特征调度机制;(iii) 改进的聚类池化策略。这些改进显著提升了分类性能与可视化可解释性。在CIFAR-100和Mini-ImageNet上的实验表明,CLUENet优于现有聚类方法与主流视觉模型,在准确率、效率与透明性之间取得出色平衡。

原文摘要 · Abstract (English)

Despite the success of convolution- and attention-based models in vision tasks, their rigid receptive fields and complex architectures limit their ability to model irregular spatial patterns and hinder interpretability, therefore posing challenges for tasks requiring high model transparency. Clustering paradigms offer promising interpretability and flexible semantic modeling, but suffer from limited accuracy, low efficiency, and gradient vanishing during training. To address these issues, we propose CLUster attEntion Network (CLUENet), an transparent deep architecture for visual semantic understanding. We propose three key innovations include (i) a Global Soft Aggregation and Hard Assignment with a Temperature-Scaled Cosin Attention and gated residual connections for enhanced local modeling, (ii) inter-block Hard and Shared Feature Dispatching, and (iii) an improved cluster pooling strategy. These enhancements significantly improve both classification performance and visual interpretability. Experiments on CIFAR-100 and Mini-ImageNet demonstrate that CLUENet outperforms existing clustering methods and mainstream visual models, offering a compelling balance of accuracy, efficiency, and transparency.

聚类注意力可解释性视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。