发现强化学习模型在渐变干扰下存在感知临界点,决定能否察觉异常。
The Boiling Frog Threshold: Criticality and Blindness in World Model-Based Anomaly Detection Under Gradual Drift
- 通过四类环境与多种检测器,发现异常检测有普遍存在的临界阈值ε*。
- 正弦漂移完全无法被任何检测器识别,体现世界模型本质特性。
- 临界点受噪声结构、检测器与环境动态三者交互影响,适合关注自监控的开发者。
当强化学习代理的观测值逐渐被破坏时,其在何种漂移速率下会‘觉醒’?我们研究了在四个MuJoCo环境中,基于世界模型的自我监控在持续观测漂移下的表现,涵盖三种检测器类型(z-score、方差、分位数)和三种模型容量。结果发现:(1)存在一个普遍的尖锐检测阈值ε*:低于该值,漂移被吸收为正常波动;高于该值,检测迅速触发。该阈值的存在及S型曲线在所有检测器类型和模型容量下均不变,但其位置取决于检测器敏感度、噪声底噪结构与环境动态的交互。(2)正弦漂移对所有检测器(包括无时间平滑的方差与分位数检测器)完全不可见,表明这是世界模型的固有属性而非检测器缺陷。(3)在每个环境中,ε*与检测器参数呈幂律关系(R²=0.89–0.97),但跨环境预测失败(R²=0.45),揭示缺失变量为环境特异性动态结构∂PE/∂ε。(4)在脆弱环境中,代理在任何检测器触发前已崩溃(‘未觉先溃’),形成根本不可监测的故障模式。研究将ε*从涌现属性重构为噪声底噪、检测器与环境动态三者之间的相互作用,为强化学习代理自监控边界提供了更可靠且实证支持的解释。
原文摘要 · Abstract (English)
When an RL agent's observations are gradually corrupted, at what drift rate does it "wake up" -- and what determines this boundary? We study world model-based self-monitoring under continuous observation drift across four MuJoCo environments, three detector families (z-score, variance, percentile), and three model capacities. We find that (1) a sharp detection threshold $\varepsilon^*$ exists universally: below it, drift is absorbed as normal variation; above it, detection occurs rapidly. The threshold's existence and sigmoid shape are invariant across all detector families and model capacities, though its position depends on the interaction between detector sensitivity, noise floor structure, and environment dynamics. (2) Sinusoidal drift is completely undetectable by all detector families -- including variance and percentile detectors with no temporal smoothing -- establishing this as a world model property rather than a detector artifact. (3) Within each environment, $\varepsilon^*$ follows a power law in detector parameters ($R^2 = 0.89$-$0.97$), but cross-environment prediction fails ($R^2 = 0.45$), revealing that the missing variable is environment-specific dynamics structure $\partial \mathrm{PE}/\partial\varepsilon$. (4) In fragile environments, agents collapse before any detector can fire ("collapse before awareness"), creating a fundamentally unmonitorable failure mode. Our results reframe $\varepsilon^*$ from an emergent world model property to a three-way interaction between noise floor, detector, and environment dynamics, providing a more defensible and empirically grounded account of self-monitoring boundaries in RL agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。