优化稳定不等于学对了,偏见会让模型稳定地错下去
Stable but Wrong: When Learning Stabilizes Away from the Truth
- 用强凸模型证明更新方向有偏会偏离真实最优解
- 多场景实验发现优化看似正常但结果持续偏离目标
- 干预后仍可修正,说明错误状态可被改变
稳定性常被视为学习成功的标志,但其反映的是优化行为而非与外部目标的正确性。本文定义‘稳定但错误’(SBW):学习过程在任务标准下保持稳定,但结果系统性偏离独立定义的目标。一个最小强凸模型表明,更新方向存在持续偏差时,收敛点会远离真实最优。在强化学习、监督学习及大语言模型持续微调中,控制实验揭示了优化看似正常与结果正确性之间的普遍分离,无论是在静态偏差还是反馈耦合偏差下均如此。恢复阶段使用干净数据和探索干预显示后续轨迹仍可被修改,但干预方式不同,未体现统一机制。结果表明,在存在持续性和反馈耦合偏差的学习系统中,优化稳定性无法作为可靠性的有效信号。
原文摘要 · Abstract (English)
Stable training is often treated as evidence that learning is succeeding, but stability characterizes optimization behavior rather than correctness relative to an external objective. We study what happens when the signal being optimized remains persistently biased. We define Stable but Wrong (SBW) as a learning state in which the learning process remains stable under a task-appropriate operational criterion while the learned outcome remains systematically displaced from an independently defined objective. A minimal strongly convex model shows that a persistent bias in the update direction can shift the unique convergence point away from the true optimum. Controlled experiments in reinforcement learning, supervised learning, and continual fine-tuning of a large language model reveal a recurring separation between apparently normal optimization and correctness under static and feedback-coupled biases. Recovery-stage clean-data access and exploration interventions further show that subsequent trajectories can remain modifiable, although the interventions differ in protocol and do not imply a shared mechanism. The results expose a basic limit of optimization stability as a reliability signal in persistent and feedback-coupled learning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。