arXiv:2412.16168cs.LGcs.CV2024-12

用主动学习探索神经元共存现象,发现其难以通过常规方法解码。

Superposition through Active Learning lens

  • 对比基线与主动学习模型在图像数据上的特征分离效果。
  • 主动学习模型在准确率和特征聚类上未显著优于基线。
  • 揭示共存现象或受样本选择偏差影响,需更精细方法应对。

超位置(Superposition)或神经元多义性是可解释性领域的重要概念,常被视为破解机器学习黑箱的复杂障碍。本文探究能否利用主动学习方法解码超位置现象。基于CIFAR-10和Tiny ImageNet数据集,使用ResNet18模型,对比基线与主动学习模型在多个评估指标下的表现,包括t-SNE可视化、余弦相似度直方图、轮廓系数(Silhouette Scores)和Davies-Bouldin指数。结果表明,主动学习模型在特征分离和整体准确率上未显著优于基线。这暗示非信息性样本选择及对不确定样本的过拟合可能阻碍了主动学习模型的泛化能力,提示解码超位置需更复杂的策略。

原文摘要 · Abstract (English)

Superposition or Neuron Polysemanticity are important concepts in the field of interpretability and one might say they are these most intricately beautiful blockers in our path of decoding the Machine Learning black-box. The idea behind this paper is to examine whether it is possible to decode Superposition using Active Learning methods. While it seems that Superposition is an attempt to arrange more features in smaller space to better utilize the limited resources, it might be worth inspecting if Superposition is dependent on any other factors. This paper uses CIFAR-10 and Tiny ImageNet image datasets and the ResNet18 model and compares Baseline and Active Learning models and the presence of Superposition in them is inspected across multiple criteria, including t-SNE visualizations, cosine similarity histograms, Silhouette Scores, and Davies-Bouldin Indexes. Contrary to our expectations, the active learning model did not significantly outperform the baseline in terms of feature separation and overall accuracy. This suggests that non-informative sample selection and potential overfitting to uncertain samples may have hindered the active learning model's ability to generalize better suggesting more sophisticated approaches might be needed to decode superposition and potentially reduce it.

可解释性主动学习超位置

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。