对抗性观测下提升状态空间模型的鲁棒性,保障强化学习决策可靠性。
Adversarial observations in probabilistic State-Space Models for robust Reinforcement Learning

- 设计对抗性观测攻击,保持观测统计一致性
- 提出方向性协方差自适应在线防御机制
- 适用于机器人等安全关键场景的抗干扰系统
在部分可观测或对抗性观测条件下,准确推断环境潜在状态及其不确定性至关重要。本文分析了对线性状态空间模型的对抗攻击,攻击者在保持观测似然约束的前提下修改观测值,确保扰动与观测模型统计一致。研究揭示此类看似合理却具有破坏性的观测如何改变潜在状态的推断,并影响下游决策及强化学习智能体性能。为此,提出一种基于方向性协方差自适应的在线贝叶斯防御方法:通过比较观测估计影响与最破坏性方向的影响,选择性削弱干扰观测的影响,同时保留正交子空间中的有效信息。该框架为构建稳健的推理与决策系统提供了理论基础,直接适用于机器人等安全关键领域,在传感器噪声、部分失效及对抗环境下仍能可靠运行。
原文摘要 · Abstract (English)
Decision-making under partial or adversarial observability requires accurate inference of the environment's latent state and its associated uncertainty. This work analyses adversarial attacks on linear state-space models, where the attacker alters observations subject to likelihood constraints that ensure that the perturbations remain statistically consistent with the observation model. We analyse how such adversarial yet plausible observations shift inference about latent states and affect downstream decision-making and the performance of reinforcement learning agents. In addition, we introduce an online Bayesian defence based on directional covariance adaptation, which selectively reduces the influence of observations by comparing their estimated impact with that of the computed most disruptive direction, while preserving information in the remaining orthogonal observation subspace. The proposed framework provides a principled approach to constructing robust inference and decision-making systems, with direct relevance to safety-critical applications such as robotics, where reliable operation under sensor noise, partial failures, and adversarial conditions is essential.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。