用内在认知缺口自动生成注意力优先级,无需外部奖励
Telogenesis: Goal Is All U Need
- 基于无知、意外和过时三类认知缺口自动生成观察目标
- 在高维环境中优先分配注意力显著提升检测效率,优势随维度上升
- 系统能自主发现环境变化规律,适合构建自适应智能体
目标导向系统通常依赖外部提供目标。我们探究注意力优先级能否从智能体内部认知状态中内生产生。提出一个优先级函数,根据三种认知缺口生成观察目标:无知(后验方差)、意外(预测误差)和过时(未观测变量置信度的时间衰减)。在两个系统中验证:最小注意力分配环境(2,000次运行)和模块化部分可观测世界(500次运行)。消融实验表明每个组件均必要。关键发现为度量依赖性反转:在全局预测误差下,覆盖式旋转更优;在变化检测延迟下,优先级引导分配胜出,且优势随维度单调增长(d = -0.95,N=48,p < 10^-6)。检测延迟服从幂律,优先级引导的指数更陡(0.55 vs. 0.40)。当衰减率对每变量可学习时,系统无需监督即可自发恢复环境波动结构(t = 22.5,p < 10^-6)。证明仅凭认知缺口,无需外部奖励,即可生成自适应优先级,优于固定策略并恢复潜在环境结构。
原文摘要 · Abstract (English)
Goal-conditioned systems assume goals are provided externally. We ask whether attentional priorities can emerge endogenously from an agent's internal cognitive state. We propose a priority function that generates observation targets from three epistemic gaps: ignorance (posterior variance), surprise (prediction error), and staleness (temporal decay of confidence in unobserved variables). We validate this in two systems: a minimal attention-allocation environment (2,000 runs) and a modular, partially observable world (500 runs). Ablation shows each component is necessary. A key finding is metric-dependent reversal: under global prediction error, coverage-based rotation wins; under change detection latency, priority-guided allocation wins, with advantage growing monotonically with dimensionality (d = -0.95 at N=48, p < 10^-6). Detection latency follows a power law in attention budget, with a steeper exponent for priority-guided allocation (0.55 vs. 0.40). When the decay rate is made learnable per variable, the system spontaneously recovers environmental volatility structure without supervision (t = 22.5, p < 10^-6). We demonstrate that epistemic gaps alone, without external reward, suffice to generate adaptive priorities that outperform fixed strategies and recover latent environmental structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。