提出新框架分析神经模型中概念表示的完整与独立性。
A framework for analyzing concept representations in neural models
- 从包含性与解耦性两维度统一评估概念子空间
- LEACE方法在两项测试中表现良好但泛化能力不足
- 语音模型中音素信息可被完整且独立表示
理解神经模型如何表征人类可解释的概念颇具挑战。现有研究从探测和概念擦除等角度探索线性概念子空间。本文提出一个统一框架,沿两个维度——包含性(概念是否仅存在于子空间内)和解耦性(是否与其他概念隔离)——分析这些子空间。在文本与语音模型上的实验表明:概念子空间可能不唯一,影响分析结果;比较五种不同社区提出的估计器发现:(1)估计器选择显著影响包含性与解耦性;(2)当前最优的概念擦除方法LEACE在两项测试中表现良好,但仍难以泛化至未见数据;(3)在HuBERT语音表示中,音素信息既被完整包含又与说话人信息解耦,而说话人信息虽能解耦却难以在紧凑子空间中实现完整包含。
原文摘要 · Abstract (English)
Understanding how neural models represent human-interpretable concepts is challenging. Prior work has explored linear concept subspaces from diverse perspectives, such as probing and concept erasure. We introduce a unified framework to study these subspaces along two axes: \textit{containment}, which tests if a concept is fully represented in a subspace but not outside it, and \textit{disentanglement}, which tests for isolation from other concepts. In experiments on both text and speech models, we first highlight that concept subspaces may not be uniquely determined, and discuss the implications for concept subspace analysis. Then, we compare properties of concept subspaces estimated using five estimators, proposed in different communities. We find that (1) the choice of estimator impacts the containment and disentanglement properties; (2) the state-of-the-art concept erasure method, LEACE, performs well on both testing axes, but still struggles to generalize to unseen data; and (3) in HuBERT speech representations, phone information is both contained and disentangled from speaker information, while speaker information is hard to contain in a compact subspace, despite being disentangled from phones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。