提出自我诊断框架,让模型学会判断自身经验的可信度。
Learning to Trust Experience: A Monitor-Trust-Regulator Framework for Learning under Unobservable Feedback Reliability
- 引入监控-信任-调节器结构,通过内部动态推断经验可信度。
- 在强化学习中实现受控怀疑,修复被错误奖励误导的系统。
- 揭示性能恢复不等于认知恢复,适合自适应学习系统设计者。
在不可观测反馈可靠性下学习面临新挑战:系统不仅需稳定学习,还需判断是否应采纳某次经验。本文研究该问题为不可观测可靠性下的知识可识别性(EIUR),其中每次经验有隐含可信度,可靠与不可靠反馈在局部无法区分,且数据由学习者自身信念和行为闭环生成。标准鲁棒学习可能稳定收敛,却形成高置信度但系统性错误的认知。本文提出元认知调控机制:一个内省控制环,从学习者内部动态中推断经验可信度。形式化为模块化监测-信任-调节器(MTR)框架,并以自我诊断为例实现——维护缓慢变化的经验信任变量,软性调节学习更新,无需外部标签或显式污染模型。实验表明,在所研究的EIUR场景中,自我诊断显著提升知识可识别性。在强化学习中,它实现校准后的怀疑态度并从系统性错误奖励中恢复;在监督学习中,暴露关键分裂现象:性能恢复不意味着认知恢复——准确率可回升,但内部信念仍被早期误导数据锁定,此失败仅能通过内省诊断发现。MTR与自我诊断共同提供内在可靠性评估的组织框架与具体设计模板。
原文摘要 · Abstract (English)
Learning under unobservable feedback reliability poses a distinct challenge beyond optimization robustness: a system must decide whether to learn from an experience, not only how to learn stably. We study this setting as Epistemic Identifiability under Unobservable Reliability (EIUR), where each experience has a latent credibility, reliable and unreliable feedback can be locally indistinguishable, and data are generated in a closed loop by the learner's own evolving beliefs and actions. In EIUR, standard robust learning can converge stably yet form high-confidence, systematically wrong beliefs. We propose metacognitive regulation as a practical response: a second, introspective control loop that infers experience credibility from endogenous evidence in the learner's internal dynamics. We formalize this as a modular Monitor-Trust-Regulator (MTR) decomposition and instantiate it with self-diagnosis, which maintains a slowly varying experience-trust variable that softly modulates learning updates, without exogenous reliability labels or an explicit corruption model. Empirically, in the EIUR regimes studied here, self-diagnosis is associated with improved epistemic identifiability. In reinforcement learning, it enables calibrated skepticism and recovery under systematically corrupted rewards. In supervised learning, it exposes a critical dissociation: performance recovery does not imply epistemic recovery. Accuracy can rebound while internal belief dynamics remain locked-in by early misleading data, a failure detectable only through introspective diagnostics. Together, MTR and self-diagnosis provide an organizing abstraction and a concrete design template for intrinsic reliability assessment in autonomous learning under unobservable reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。