arXiv:2601.11625cs.AIcs.LG2026-01

通过追踪注意力变化,找出模型决策依据稳定的训练阶段。

Reasoning Stabilization Point: A Training-Time Signal for Stable Evidence and Shortcut Reliance

  • 用词级归因变化衡量微调过程中的推理稳定性。
  • 多数任务在训练早期就进入低漂移稳定区,准确率仍可提升。
  • 能发现模型悄悄依赖错误线索,适合关注可靠性的研究者。

微调预训练语言模型虽可提升任务性能,但会悄然改变模型依赖的证据。本文提出一种训练时可解释性视角,跟踪微调各轮次中词级归因的变化。定义‘解释漂移’为固定测试集上归一化词归因在不同轮次间的变动,并引入‘推理稳定点’(RSP),即漂移持续保持低位的最早训练轮次。RSP基于训练内部漂移动态计算,无需在分布外数据上调参。在多个轻量级Transformer分类器和基准分类任务中,漂移通常在训练初期即进入低且稳定的区间,而验证准确率仍略有波动。在含标签相关触发词的控制实验中,归因动态显示模型对捷径的依赖不断加剧,即使验证准确率仍表现良好。总体而言,解释漂移提供了一种简单、低成本的诊断工具,可用于监控微调过程中决策证据的演变,并选取处于稳定证据区的检查点。

原文摘要 · Abstract (English)

Fine-tuning pretrained language models can improve task performance while subtly altering the evidence a model relies on. We propose a training-time interpretability view that tracks token-level attributions across finetuning epochs. We define explanation driftas the epoch-to-epoch change in normalized token attributions on a fixed probe set, and introduce the Reasoning Stabilization Point(RSP), the earliest epoch after which drift remains consistently low. RSP is computed from within-run drift dynamics and requires no tuning on out-of-distribution data. Across multiple lightweight transformer classifiers and benchmark classification tasks, drift typically collapses into a low, stable regime early in training, while validation accuracy continues to change only marginally. In a controlled shortcut setting with label-correlated trigger tokens, attribution dynamics expose increasing reliance on the shortcut even when validation accuracy remains competitive. Overall, explanation drift provides a simple, low-cost diagnostic for monitoring how decision evidence evolves during fine-tuning and for selecting checkpoints in a stable-evidence regime.

模型可解释性微调优化推理稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。