arXiv:2412.02412cs.LGcs.AI2024-12

用二维语义空间可视化神经网络内部表示,揭示隐藏模式。

VISTA: A Panoramic View of Neural Representations

  • 将高维神经表示映射到可读的二维语义空间
  • 在稀疏自编码器潜空间中发现新特征与关联模式
  • 适合研究模型可解释性与内部表征的科研人员

我们提出VISTA(内部状态及其关联的可视化),一种用于视觉探索和解释神经网络表示的新流程。VISTA通过将现代机器学习模型中的庞大多维空间映射到语义2D空间,解决了分析复杂表示的挑战。生成的拼贴图直观展现内部表示中的模式与关系。我们通过应用于稀疏自编码器潜空间,揭示了新的属性与解读。本文介绍VISTA方法,展示案例研究结果(https://got.drib.net/latents/),并讨论其对机器学习多个领域中神经网络可解释性的意义。

原文摘要 · Abstract (English)

We present VISTA (Visualization of Internal States and Their Associations), a novel pipeline for visually exploring and interpreting neural network representations. VISTA addresses the challenge of analyzing vast multidimensional spaces in modern machine learning models by mapping representations into a semantic 2D space. The resulting collages visually reveal patterns and relationships within internal representations. We demonstrate VISTA's utility by applying it to sparse autoencoder latents uncovering new properties and interpretations. We review the VISTA methodology, present findings from our case study ( https://got.drib.net/latents/ ), and discuss implications for neural network interpretability across various domains of machine learning.

可解释性神经网络可视化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。