用信息论指标追踪强化学习中的策略一致性,可精准诊断传感器与执行器故障。
Mutual Information Tracks Policy Coherence in Reinforcement Learning
- 通过状态-动作互信息变化揭示学习过程中的注意力演化规律。
- 状态-动作互信息从0.84增至2.83比特(增长238%),而状态熵上升仍保持高关联。
- 不同故障类型在信息指标上呈现可区分特征,支持无损故障定位。
部署于真实环境的强化学习(RL)智能体易受传感器故障、执行器磨损和环境漂移影响,但缺乏内在的故障检测与诊断机制。本文提出一种信息论框架,揭示了强化学习的基本动态,并提供实用的部署期异常诊断方法。通过对机器人控制任务中状态-动作互信息的分析,我们首次证明:成功学习具有特定的信息特征——尽管状态熵上升,状态与动作间的互信息仍从0.84比特稳步增长至2.83比特(增长238%),表明智能体逐步聚焦于任务相关模式。令人关注的是,状态、动作与下一状态的联合互信息MI(S,A;S')呈倒U型曲线,早期学习阶段达到峰值后下降,反映出从广泛探索向高效利用的转变。更实际地,我们发现信息度量可差异化诊断系统故障:观测空间噪声(传感器故障)导致所有信息通道普遍崩溃,显著降低状态-动作耦合;而动作空间噪声(执行器故障)仅破坏动作-结果可预测性,却保留状态-动作关系。这一差异性诊断能力通过受控扰动实验验证,实现精确故障定位,无需架构修改或性能损失。本研究将信息模式确立为学习状态与系统健康状况的双重标志,为具备自主故障检测与策略调整能力的自适应强化学习系统奠定基础。
原文摘要 · Abstract (English)
Reinforcement Learning (RL) agents deployed in real-world environments face degradation from sensor faults, actuator wear, and environmental shifts, yet lack intrinsic mechanisms to detect and diagnose these failures. We present an information-theoretic framework that reveals both the fundamental dynamics of RL and provides practical methods for diagnosing deployment-time anomalies. Through analysis of state-action mutual information patterns in a robotic control task, we first demonstrate that successful learning exhibits characteristic information signatures: mutual information between states and actions steadily increases from 0.84 to 2.83 bits (238% growth) despite growing state entropy, indicating that agents develop increasingly selective attention to task-relevant patterns. Intriguingly, states, actions and next states joint mutual information, MI(S,A;S'), follows an inverted U-curve, peaking during early learning before declining as the agent specializes suggesting a transition from broad exploration to efficient exploitation. More immediately actionable, we show that information metrics can differentially diagnose system failures: observation-space, i.e., states noise (sensor faults) produces broad collapses across all information channels with pronounced drops in state-action coupling, while action-space noise (actuator faults) selectively disrupts action-outcome predictability while preserving state-action relationships. This differential diagnostic capability demonstrated through controlled perturbation experiments enables precise fault localization without architectural modifications or performance degradation. By establishing information patterns as both signatures of learning and diagnostic for system health, we provide the foundation for adaptive RL systems capable of autonomous fault detection and policy adjustment based on information-theoretic principles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。