arXiv:2505.06279cs.LGstat.ML2025-05

解析无监督强化学习中好奇心如何影响注意力与行为

Interpretable Learning Dynamics in Unsupervised Reinforcement Learning

  • 用注意力分析与动态指标揭示内在动机对代理的影响
  • Transformer-RND展现最广视野、最高探索率和紧凑表征
  • 适合关注智能体可解释性与内在动机机制的研究者

我们提出一个无监督强化学习(URL)智能体的可解释性框架,旨在理解内在动机如何塑造注意力、行为与表征学习。分析了在程序化生成环境上训练的五种智能体:DQN、RND、ICM、PPO 和 Transformer-RND 变体,采用 Grad-CAM、层间相关传播(LRP)、探索度量和潜在空间聚类方法。为捕捉智能体随时间的感知与适应过程,引入两个新指标:注意力多样性(衡量关注的空间范围)和注意力变化率(量化注意力的时间变动)。结果表明,基于好奇心的智能体相比外在动机模型展现出更宽泛、更动态的注意力与探索行为。其中,Transformer-RND 结合了广泛注意力、高探索覆盖率和紧凑结构化的潜在表征。研究凸显了架构先验偏置与训练信号对内部动态的影响。该框架超越传统奖励评估,提供诊断工具以探究强化学习智能体的感知与抽象机制,推动更可解释、泛化性更强的行为设计。

原文摘要 · Abstract (English)

We present an interpretability framework for unsupervised reinforcement learning (URL) agents, aimed at understanding how intrinsic motivation shapes attention, behavior, and representation learning. We analyze five agents DQN, RND, ICM, PPO, and a Transformer-RND variant trained on procedurally generated environments, using Grad-CAM, Layer-wise Relevance Propagation (LRP), exploration metrics, and latent space clustering. To capture how agents perceive and adapt over time, we introduce two metrics: attention diversity, which measures the spatial breadth of focus, and attention change rate, which quantifies temporal shifts in attention. Our findings show that curiosity-driven agents display broader, more dynamic attention and exploratory behavior than their extrinsically motivated counterparts. Among them, TransformerRND combines wide attention, high exploration coverage, and compact, structured latent representations. Our results highlight the influence of architectural inductive biases and training signals on internal agent dynamics. Beyond reward-centric evaluation, the proposed framework offers diagnostic tools to probe perception and abstraction in RL agents, enabling more interpretable and generalizable behavior.

强化学习可解释性好奇心驱动注意力分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。