提出主动感知框架,破解状态与动态不确定性耦合难题。
Breaking the Epistemic Trap: Active Perception Under Compound Uncertainty

- 用互信息量化状态与动态不确定性的耦合程度
- 信息探索策略使系统主动探测环境,提升认知能力
- 安全约束随不确定性自动收紧,适合高风险决策场景
在自动驾驶、医疗决策等安全关键领域部署强化学习时,系统常因遭遇未知条件而失效。本文指出根本瓶颈并非单一挑战(如动态变化或观测不全),而是它们的协同作用——即“认知陷阱”:状态估计依赖动态知识,而动态学习又需准确状态信息。模拟运动实验显示,复合不确定性导致的性能下降达77%,远超各因素单独影响之和(46%),表明耦合失败模式会显著加剧。传统方法多为被动认知,无法解决此问题。本文提出将安全性重构为信息问题,构建自适应安全架构:首先引入复合不确定性系数κ,基于互信息衡量耦合强度;其次设计最大信息强化学习(MaxInfoRL)策略,主动探测系统动态;最后实现随认知耦合增强而自动收紧的安全约束。三者结合,推动从被动鲁棒性到主动感知的范式转变,为不确定环境下决策系统提供可解释、可自省、可行动的认知路径。
原文摘要 · Abstract (English)
Deploying reinforcement learning in safety critical domains, from autonomous vehicles to medical decision support, is constrained by failures arising when systems encounter unfamiliar conditions. We argue that the fundamental bottleneck is not individual challenges like changing dynamics or incomplete observations, but their synergistic interaction, which we term the Epistemic Trap: agents cannot estimate their state without knowing system dynamics, nor learn dynamics without accurate state information. Proof-of-concept experiments in simulated locomotion reveal that combining these uncertainties causes failures far worse than either challenge alone, a 77% observed degradation against the 46% additive prediction, demonstrating that compounding failure modes can emerge and, when they do, far exceed what additive reasoning would predict. Conventional approaches typically adopt a passive epistemic stance that cannot resolve this coupled uncertainty. We propose reframing safety as an information problem. We introduce an Adaptive Safety Architecture built around three contributions. First, the Compound Uncertainty Coefficient ($κ$), a mutual-information based metric that quantifies how tightly state and dynamics uncertainties are coupled. Second, information-seeking policies governed by a MaxInfoRL objective that actively probe system dynamics rather than waiting for the environment to reveal itself passively. Third, regime adaptive safety constraints that tighten automatically as epistemic coupling rises. Together, these constitute a paradigm shift from passive robustness to active perception, offering a principled path toward decision making systems that operate under uncertainty, recognize their own ignorance, and act strategically to resolve it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。