固定解释数据也能让模型生成真实自我反思,无需实时更新监督信号。
Introspective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision

- 用模型自身或同类模型的旧解释作为监督信号训练新解释。
- 解释与当前行为相关性高时,能准确追踪行为变化。
- 适用于拒答、奉承等任务,对标签噪声鲁棒,适合大规模后训练。
语言模型在生成预测解释时,何时能实现真实自我反思而非表面模仿?我们研究了使用模型在修改输入后的反事实行为作为监督信号,训练其解释影响自身决策的输入特征。令人惊讶的是,即使使用早期检查点或不同模型家族中行为相似模型的固定反事实解释,模型仍能生成更符合自身当前行为的解释。这种“内省耦合”现象出现在训练过程中解释与当前行为保持足够相关性时,即便行为本身已发生变化。我们还发现,当解释训练与其他后训练目标并行进行时,解释能自动追踪行为变化,无需更新监督信号。该现象在多个任务(如奉承和拒绝)中均出现,且对标签噪声具有鲁棒性。结果表明,即使使用固定的反事实解释数据集,也能为内省提供可扩展、通用的后训练信号。
原文摘要 · Abstract (English)
When does training language models (LMs) to generate explanations of their predictions yield faithful introspection, rather than superficial imitation? We study LMs trained to explain which features of their inputs influenced their behavior, using models' counterfactual behavior on modified inputs as supervision. Surprisingly, we find that LMs trained on fixed counterfactual explanations derived from earlier checkpoints of themselves, or even from behaviorally similar models in different families, frequently produce explanations more faithful to their own current behaviors than to those of their training targets. This "introspective" coupling between LM explanations and behaviors occurs when training explanations remain sufficiently correlated with current behaviors over the course of training, even as behaviors themselves shift. We also show that introspective coupling tracks behavior shifts: when explanation training is provided concurrently with other post-training objectives, explanations track those shifts without requiring updated supervision. This phenomenon appears in multiple tasks, including sycophancy and refusal, and is robust to label noise. Overall, our results show that even fixed datasets of counterfactual explanations can provide scalable and generalizable post-training signal for introspection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。