用贝叶斯方法解析神经网络任务表征的因果贡献。
Understanding Task Representations in Neural Networks via Bayesian Ablation

- 基于贝叶斯框架建模表征单元的分布,推断其对任务的贡献
- 提出信息论工具,量化表征的分布式程度与多义性
- 适合研究模型可解释性与认知建模的学者参考
神经网络因其灵活性和涌现特性,是认知建模的强大工具。然而,由于其表征具有非符号语义,理解其学习到的表示仍具挑战性。本文提出一种新颖的概率框架,用于解析神经网络中的潜在任务表征。受贝叶斯推理启发,该方法在表征单元上定义分布,以推断其对任务性能的因果贡献。结合信息论思想,我们提出一套工具与度量,揭示模型的关键属性,包括表征的分布式程度、流形复杂性及多义性。
原文摘要 · Abstract (English)
Neural networks are powerful tools for cognitive modeling due to their flexibility and emergent properties. However, interpreting their learned representations remains challenging due to their sub-symbolic semantics. In this work, we introduce a novel probabilistic framework for interpreting latent task representations in neural networks. Inspired by Bayesian inference, our approach defines a distribution over representational units to infer their causal contributions to task performance. Using ideas from information theory, we propose a suite of tools and metrics to illuminate key model properties, including representational distributedness, manifold complexity, and polysemanticity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。