通过信息瓶颈压缩多传感器数据,提升强化学习的样本效率与鲁棒性。
Multimodal Information Bottleneck for Deep Reinforcement Learning with Multiple Sensors
- 设计多模态信息瓶颈模型,压缩无关信息保留任务相关特征。
- 在多种运动任务中实现更优采样效率和零样本抗噪声能力。
- 适合需要融合视觉与本体感知的机器人控制场景。
强化学习在机器人控制任务中表现优异,但在利用多感官模态(如第一人称图像与本体感知)时面临信息利用效率低的问题。现有方法通过重构或互信息构建辅助损失来提取联合表征,但可能引入与策略学习无关的信息,降低性能。本文提出多模态信息瓶颈模型,对原始多模态观测进行压缩,保留对未来动作预测有帮助的信息,同时过滤无关内容。该模型通过最小化上界目标实现可计算优化。在多个包含第一人称图像与本体感知的挑战性运动任务上,实验表明该方法相比主流基线具有更高样本效率,并在未见过的白噪声干扰下展现更强零样本鲁棒性。实证还证明,结合视觉与本体感知比单一模态更有助于学习运动策略。
原文摘要 · Abstract (English)
Reinforcement learning has achieved promising results on robotic control tasks but struggles to leverage information effectively from multiple sensory modalities that differ in many characteristics. Recent works construct auxiliary losses based on reconstruction or mutual information to extract joint representations from multiple sensory inputs to improve the sample efficiency and performance of reinforcement learning algorithms. However, the representations learned by these methods could capture information irrelevant to learning a policy and may degrade the performance. We argue that compressing information in the learned joint representations about raw multimodal observations is helpful, and propose a multimodal information bottleneck model to learn task-relevant joint representations from egocentric images and proprioception. Our model compresses and retains the predictive information in multimodal observations for learning a compressed joint representation, which fuses complementary information from visual and proprioceptive feedback and meanwhile filters out task-irrelevant information in raw multimodal observations. We propose to minimize the upper bound of our multimodal information bottleneck objective for computationally tractable optimization. Experimental evaluations on several challenging locomotion tasks with egocentric images and proprioception show that our method achieves better sample efficiency and zero-shot robustness to unseen white noise than leading baselines. We also empirically demonstrate that leveraging information from egocentric images and proprioception is more helpful for learning policies on locomotion tasks than solely using one single modality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。