arXiv:2604.11914cs.AI2026-04

将自监控模块整合到决策路径中,可提升智能体在动态环境中的表现。

Self-Monitoring Benefits from Structural Integration: Lessons from Metacognition in Continuous-Time Multi-Timescale Agents

  • 通过结构化整合自信度、意外感等自监控信号来指导探索与策略
  • 整合后在非平稳环境中性能提升显著(Cohen's d = 0.62)
  • 自监控需嵌入决策流程,而非作为附加模块使用

自监控能力——元认知、自我预测和主观时长——常被视为强化学习智能体的有益补充。我们在连续时间多时标智能体上研究了这一问题,其在复杂度各异的捕食者-猎物生存环境中运行,包括一个二维部分可观测变体。实验表明,三种以辅助损失形式添加的自监控模块,在20个随机种子、一维与二维环境及训练步数达50,000的情况下,均未带来统计显著的性能提升。诊断发现,这些模块输出趋近恒定(置信度标准差 < 0.006,注意力分配标准差 < 0.011),主观时长机制对折扣因子影响不足0.03%。策略敏感性分析显示,智能体决策不受模块输出影响。进一步发现,将模块输出结构化整合——用自信度控制探索、意外感触发工作区广播、自模型预测作为策略输入——在非平稳环境中带来中等至较大改进(Cohen's d = 0.62,p = 0.06,配对检验)。组件消融显示,从时标记忆到策略的通路贡献最大。但该整合方法并未显著优于无自监控基线(d = 0.15,p = 0.67),且参数匹配的无模块对照组表现相当,说明收益可能来自恢复被忽略模块带来的趋势性损害,而非自监控内容本身。架构启示是:自监控应置于决策路径上,而非旁侧。

原文摘要 · Abstract (English)

Self-monitoring capabilities -- metacognition, self-prediction, and subjective duration -- are often proposed as useful additions to reinforcement learning agents. But do they actually help? We investigate this question in a continuous-time multi-timescale agent operating in predator-prey survival environments of varying complexity, including a 2D partially observable variant. We first show that three self-monitoring modules, implemented as auxiliary-loss add-ons to a multi-timescale cortical hierarchy, provide no statistically significant benefit across 20 random seeds, 1D and 2D predator-prey environments with standard and non-stationary variants, and training horizons up to 50,000 steps. Diagnosing the failure, we find the modules collapse to near-constant outputs (confidence std < 0.006, attention allocation std < 0.011) and the subjective duration mechanism shifts the discount factor by less than 0.03%. Policy sensitivity analysis confirms the agent's decisions are unaffected by module outputs in this design. We then show that structurally integrating the module outputs -- using confidence to gate exploration, surprise to trigger workspace broadcasts, and self-model predictions as policy input -- produces a medium-large improvement over the add-on approach (Cohen's d = 0.62, p = 0.06, paired) in a non-stationary environment. Component-wise ablations reveal that the TSM-to-policy pathway contributes most of this gain. However, structural integration does not significantly outperform a baseline with no self-monitoring (d = 0.15, p = 0.67), and a parameter-matched control without modules performs comparably, so the benefit may lie in recovering from the trend-level harm of ignored modules rather than in self-monitoring content. The architectural implication is that self-monitoring should sit on the decision pathway, not beside it.

强化学习自监控元认知决策系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。