通过扰动训练揭示语言模型中的表征学习机制。
Perturbation: A simple and efficient adversarial tracer for representation learning in language models
- 用单个对抗样本微调模型,观察扰动传播范围
- 在不同语言粒度上发现结构化信息迁移现象
- 无需几何假设,适用于已训练模型且不误判未训练模型
深度神经语言模型中的语言表征学习已研究数十年,但如何识别模型内部表征仍是未解难题。传统方法或强加不合理的几何约束(如线性),或使表征概念过于空泛。本文提出新视角:将表征视为学习的通道而非激活模式。方法简单:对语言模型进行单个对抗样本微调,测量该扰动对其他样本的影响范围。该方法不依赖几何假设,且不会在未训练模型中误检表征。在已训练模型中,扰动可揭示多层次语言粒度上的结构化转移,表明语言模型既沿表征路径泛化,又能仅通过经验习得语言抽象。结果支持模型具备内在表征组织能力。
原文摘要 · Abstract (English)
Linguistic representation learning in deep neural language models (LMs) has been studied for decades, for both practical and theoretical reasons. However, finding representations in LMs remains an unsolved problem, in part due to a dilemma between enforcing implausible constraints on representations (e.g., linearity; Arora et al. 2024) and trivializing the notion of representation altogether (Sutter et al., 2025). Here we escape this dilemma by reconceptualizing representations not as patterns of activation but as conduits for learning. Our approach is simple: we perturb an LM by fine-tuning it on a single adversarial example and measure how this perturbation ``infects'' other examples. Perturbation makes no geometric assumptions, and unlike other methods, it does not find representations where it should not (e.g., in untrained LMs). But in trained LMs, perturbation reveals structured transfer at multiple linguistic grain sizes, suggesting that LMs both generalize along representational lines and acquire linguistic abstractions from experience alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。