噪声监督下学习的瓶颈源于反馈与真相的速率不匹配。
Learning under noisy supervision is governed by a feedback-truth gap
- 提出双时标模型,揭示反馈速度与任务结构评估速度不一致时必然产生偏差。
- 实验验证在2700次神经网络训练和人类学习中,偏差普遍存在且影响行为决策。
- 不同系统调节偏差方式不同:神经网络记忆偏差,人类短期过调后自我修正。
当反馈吸收速度超过任务结构评估速度时,学习者会更依赖反馈而非真相。双时标模型表明,只要两者速率不一致,反馈-真相差距就不可避免,仅当速率匹配时才消失。我们在30个数据集上进行2700次神经网络噪声标签训练、292名被试的人类概率反转学习、25名被试的含同步脑电的人类奖惩学习中验证了该预测。每个系统中,‘真相’被操作性定义为:保留标签、客观正确选项或被试反馈前的预期——唯一可从反馈后脑电信号解码的非循环参照。该差距在所有系统中普遍存在,但调控机制不同:密集网络将其积累为记忆;稀疏残差结构抑制其增长;人类则产生短暂过度投入并主动恢复。神经层面的过度投入(~0.04–0.10)被放大十倍,转化为行为上的显著承诺(d = 3.3–3.9)。这一差距是噪声监督学习的根本约束,其后果取决于各系统采用的调节策略。
原文摘要 · Abstract (English)
When feedback is absorbed faster than task structure can be evaluated, the learner will favor feedback over truth. A two-timescale model shows this feedback-truth gap is inevitable whenever the two rates differ and vanishes only when they match. We test this prediction across neural networks trained with noisy labels (30 datasets, 2,700 runs), human probabilistic reversal learning (N = 292), and human reward/punishment learning with concurrent EEG (N = 25). In each system, truth is defined operationally: held-out labels, the objectively correct option, or the participant's pre-feedback expectation - the only non-circular reference decodable from post-feedback EEG. The gap appeared universally but was regulated differently: dense networks accumulated it as memorization; sparse-residual scaffolding suppressed it; humans generated transient over-commitment that was actively recovered. Neural over-commitment (~0.04-0.10) was amplified tenfold into behavioral commitment (d = 3.3-3.9). The gap is a fundamental constraint on learning under noisy supervision; its consequences depend on the regulation each system employs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。