arXiv:2606.12360cs.LG2026-06

用可解释性分析数据,让模型学习更可控、更符合预期。

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal

论文配图:Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal
图 1 · 摘自论文原文
  • 通过可解释性工具识别偏好数据中的关键概念差异
  • 实验证明能减少模型学偏斜行为,增强安全与个性特征
  • 适合关注模型行为可控性的研究者和工程师

语言模型后训练是塑造模型行为的主要阶段,但目前仍依赖抽象的标量奖励来汇总多样目标。这种抽象使实践者难以看清数据实际教给模型的内容,导致模型学习到虚假关联,引发过度风格化和迎合等问题。为解决此问题,我们提出一种以数据为中心的后训练流程:利用可解释性协议,对偏好与非偏好生成结果间的潜在概念差异建立统计假设,使其可视化并支持细粒度用户反馈。在此基础上,我们将多种基于可解释性的训练方法统一为通过特征或数据干预来塑造奖励信号的手段。实证表明,该流程能诊断现有偏好数据中的不良信号,缓解非目标学习,并可有效放大或调控如安全机制和模型人格等期望属性。更广泛地,结果表明可解释性可将后训练从优化模糊代理奖励,转变为对学习信号本身的审计与雕琢。

原文摘要 · Abstract (English)

Language-model post-training is the main stage at which model behavior is shaped, yet it still largely involves optimization of scalar rewards that summarize diverse desiderata. This abstraction gives practitioners little visibility into what their data actually teaches models, allowing spurious correlations to be learned by a model and inducing undesirable behaviors such as over-stylization and sycophancy. To address this problem, we ask: can we inspect a preference dataset before optimization and decide, at the level of concepts, which behaviors a model should be allowed to learn? Motivated by this, we introduce a data-centric post-training pipeline that uses interpretability protocols to develop statistical hypotheses for the latent concepts separating preferred from dispreferred generations, making them explicit for fine-grained user feedback. Building on this view, we unify several interpretability-based training protocols as ways of shaping rewards via feature or data interventions. Empirically, we show that our pipeline diagnoses undesirable signals in existing preference data, mitigates off-target learning, and can also help amplify or shape desired properties such as safeguards and model personality. More broadly, our results suggest that interpretability can turn post-training from optimizing opaque proxy rewards into a process of auditing and sculpting the learning signal itself.

可解释性后训练偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。