arXiv:2504.10071cs.AI2025-04被引 2

提出可解释的视觉强化学习特征提取器,让模型注意力精准定位目标物体位置。

Pay Attention to What and Where? Interpretable Feature Extractor in Vision-based Deep Reinforcement Learning

  • 设计双模块结构:人类可读编码+代理友好编码
  • 在57个ATARI游戏上实现空间定位准确、解释性强
  • 适用于需要可解释性的强化学习场景

现有可解释深度强化学习方法存在注意力掩码与视觉输入中物体位置错位的问题。本文针对传统卷积神经网络的空间局限性,提出可解释特征提取器(IFE)架构,旨在生成精确的注意力掩码,清晰展现智能体在空间域中的关注点‘何物’与‘何处’。设计包含人类可理解编码模块以生成完全可解释的注意力掩码,以及代理友好编码模块以提升智能体学习效率。两者协同构建面向视觉强化学习的可解释特征提取器,使注意力掩码在空间维度上保持一致、对人类高度可读且有效突出视觉输入中的关键对象或区域。IFE被集成至Fast and Data-efficient Rainbow框架,在57个ATARI游戏上验证了其在空间保真度、可解释性与数据效率方面的有效性。最后,通过将其引入异步优势演员-评论家模型(A3C),展示了方法的通用性。

原文摘要 · Abstract (English)

Current approaches in Explainable Deep Reinforcement Learning have limitations in which the attention mask has a displacement with the objects in visual input. This work addresses a spatial problem within traditional Convolutional Neural Networks (CNNs). We propose the Interpretable Feature Extractor (IFE) architecture, aimed at generating an accurate attention mask to illustrate both "what" and "where" the agent concentrates on in the spatial domain. Our design incorporates a Human-Understandable Encoding module to generate a fully interpretable attention mask, followed by an Agent-Friendly Encoding module to enhance the agent's learning efficiency. These two components together form the Interpretable Feature Extractor for vision-based deep reinforcement learning to enable the model's interpretability. The resulting attention mask is consistent, highly understandable by humans, accurate in spatial dimension, and effectively highlights important objects or locations in visual input. The Interpretable Feature Extractor is integrated into the Fast and Data-efficient Rainbow framework, and evaluated on 57 ATARI games to show the effectiveness of the proposed approach on Spatial Preservation, Interpretability, and Data-efficiency. Finally, we showcase the versatility of our approach by incorporating the IFE into the Asynchronous Advantage Actor-Critic Model.

可解释AI强化学习注意力机制视觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。