arXiv:2502.01316cs.LGcs.AI2025-02ICML被引 2

多视角观测下学习任务相关状态表示,提升控制鲁棒性

Learning Fused State Representations for Control from Multi-View Observations

  • 引入双模拟度量学习,聚焦任务相关特征
  • 设计视图掩码与隐空间重建辅助任务,增强缺失视图容错能力
  • 在干扰和缺视场景中仍保持高精度,适合复杂视觉控制任务

多视角强化学习(MVRL)通过提供多视角观测,使智能体能更有效、更精准地感知环境。近年来的研究集中于从多视角观测中提取潜在表示,并用于控制任务。然而,在存在冗余、干扰信息或视角缺失的情况下,学习紧凑且任务相关的表示仍具挑战。本文提出多视角融合状态控制方法(MFSC),首次将双模拟度量学习引入MVRL,以学习任务相关表示。此外,我们设计了一种基于多视角的掩码机制与隐空间重建辅助任务,利用跨视角共享信息,通过引入掩码标记提升模型在缺失视图下的鲁棒性。大量实验表明,该方法在多种MVRL任务中优于现有方法,即使在存在干扰或视角缺失的现实场景中,性能依然稳定可靠。

原文摘要 · Abstract (English)

Multi-View Reinforcement Learning (MVRL) seeks to provide agents with multi-view observations, enabling them to perceive environment with greater effectiveness and precision. Recent advancements in MVRL focus on extracting latent representations from multiview observations and leveraging them in control tasks. However, it is not straightforward to learn compact and task-relevant representations, particularly in the presence of redundancy, distracting information, or missing views. In this paper, we propose Multi-view Fusion State for Control (MFSC), firstly incorporating bisimulation metric learning into MVRL to learn task-relevant representations. Furthermore, we propose a multiview-based mask and latent reconstruction auxiliary task that exploits shared information across views and improves MFSC's robustness in missing views by introducing a mask token. Extensive experimental results demonstrate that our method outperforms existing approaches in MVRL tasks. Even in more realistic scenarios with interference or missing views, MFSC consistently maintains high performance.

多视角强化学习状态表示鲁棒控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。