arXiv:2605.19908cs.CL2026-05

同一模型不同评分机制,性能相差4倍,关键在信号何时被整合。

Where Does Authorship Signal Emerge in Encoder-Based Language Models?

论文配图:Where Does Authorship Signal Emerge in Encoder-Based Language Models?
图 1 · 摘自论文原文
  • 通过因果干预发现,评分器决定作者特征在编码器中何时被整合。
  • 平均池化使特征早期整合,而后期交互延迟整合至深层。
  • 差异源于评分器梯度结构,导致训练路径截然不同。

使用相同预训练编码器、数据和损失函数微调的作者归属模型,仅因评分机制不同,性能可相差四倍。我们利用机制可解释性工具分析这一差距。结果显示,诸如词长、标点密度和虚词频率等风格特征在所有层中均同样存在,包括一个现成的控制编码器,表明该差距并非由线性可读性造成。因果干预显示,评分器决定了编码器何时整合作者特征信号:平均池化促使特征在早期到中期层完成整合,而后期交互则将整合推迟至深层。我们进一步从评分器的梯度结构推导出这一差异,并发现训练动态呈现不同的学习轨迹,由此产生显著性能分化。

原文摘要 · Abstract (English)

Authorship attribution models fine-tuned with the same pretrained encoder, data, and loss can differ four-fold in performance depending only on their scoring mechanism. We use mechanistic interpretability tools to explain this gap. Stylistic features such as word length, punctuation density, and function-word frequency are similarly available at every layer in every model we probe, including an off-the-shelf control encoder, suggesting that the gap is not explained by their linear readability. Instead, causal intervention shows that the scorer appears to determine where the encoder consolidates authorship signal. Mean pooling forces consolidation by early to mid layers, while late interaction defers it to later layers. We further derive this difference from the gradient structure of each scorer, and training dynamics reveal distinct learning trajectories that follow from that difference.

作者归属可解释性编码器评分机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。