用几何距离建模状态不确定性,提升强化学习的鲁棒性
Geometry of Uncertainty: Learning Metric Spaces for Multimodal State Estimation in RL
- 通过学习可度量的潜在空间,让距离反映状态间转移所需最少动作数
- 在多模态任务中显著降低对传感器噪声的敏感度,优于基线方法
- 适合需要高鲁棒性状态估计的强化学习系统设计者
从高维、多模态且含噪的观测中估计环境状态是强化学习中的核心挑战。传统方法依赖概率模型处理不确定性,但常需显式假设噪声分布,限制泛化能力。本文提出一种新方法,学习结构化的潜在表示,使状态间的距离直接对应于状态转移所需的最少动作数。该度量空间提供了无需显式概率建模的不确定性几何解释。为此,我们引入多模态潜在转移模型与基于逆距离加权的传感器融合机制,实现多模态传感器数据的自适应融合,无需事先知晓噪声分布。我们在多个多模态强化学习任务上验证了该方法,结果表明其对传感器噪声更具鲁棒性,状态估计性能优于基线方法。实验还显示,使用该表示可提升强化学习智能体表现,无需额外的噪声增强。结果表明,利用转移感知的度量空间为序列决策中的鲁棒状态估计提供了一种原理性强且可扩展的解决方案。
原文摘要 · Abstract (English)
Estimating the state of an environment from high-dimensional, multimodal, and noisy observations is a fundamental challenge in reinforcement learning (RL). Traditional approaches rely on probabilistic models to account for the uncertainty, but often require explicit noise assumptions, in turn limiting generalization. In this work, we contribute a novel method to learn a structured latent representation, in which distances between states directly correlate with the minimum number of actions required to transition between them. The proposed metric space formulation provides a geometric interpretation of uncertainty without the need for explicit probabilistic modeling. To achieve this, we introduce a multimodal latent transition model and a sensor fusion mechanism based on inverse distance weighting, allowing for the adaptive integration of multiple sensor modalities without prior knowledge of noise distributions. We empirically validate the approach on a range of multimodal RL tasks, demonstrating improved robustness to sensor noise and superior state estimation compared to baseline methods. Our experiments show enhanced performance of an RL agent via the learned representation, eliminating the need of explicit noise augmentation. The presented results suggest that leveraging transition-aware metric spaces provides a principled and scalable solution for robust state estimation in sequential decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。