解析内在动机如何改变决策变换器的表示结构,提升离线强化学习性能。
Toward Explainable Offline RL: Analyzing Representations in Intrinsically Motivated Decision Transformers
- 通过后验分析框架研究内在动机对决策变换器嵌入表示的影响。
- 发现不同动机机制生成差异化的嵌入结构,且与环境相关性能提升有关。
- 揭示内在动机是塑造表示几何的先验,而非简单探索奖励。
弹性决策变换器(EDTs)在离线强化学习中表现优异,提供了一个将序列建模与不确定性下的决策统一的灵活框架。近期研究表明,在EDTs中引入内在动机机制可提升探索任务性能,但其背后的表示机制仍不明确。本文提出一种系统性的后验可解释性分析框架,探究内在动机如何影响EDTs中的学习嵌入。通过对嵌入属性(包括协方差结构、向量范数和正交性)的统计分析,我们发现不同内在动机变体产生了根本不同的表示结构。分析表明,嵌入指标与性能之间存在环境特定的相关模式,解释了为何内在动机能改善策略学习。这些发现表明,内在动机的作用远超简单的探索奖励,而是作为表示先验,以生物上合理的方式塑造嵌入几何,形成促进更好决策的环境特定组织结构。
原文摘要 · Abstract (English)
Elastic Decision Transformers (EDTs) have proved to be particularly successful in offline reinforcement learning, offering a flexible framework that unifies sequence modeling with decision-making under uncertainty. Recent research has shown that incorporating intrinsic motivation mechanisms into EDTs improves performance across exploration tasks, yet the representational mechanisms underlying these improvements remain unexplored. In this paper, we introduce a systematic post-hoc explainability framework to analyze how intrinsic motivation shapes learned embeddings in EDTs. Through statistical analysis of embedding properties (including covariance structure, vector magnitudes, and orthogonality), we reveal that different intrinsic motivation variants create fundamentally different representational structures. Our analysis demonstrates environment-specific correlation patterns between embedding metrics and performance that explain why intrinsic motivation improves policy learning. These findings show that intrinsic motivation operates beyond simple exploration bonuses, acting as a representational prior that shapes embedding geometry in biologically plausible ways, creating environment-specific organizational structures that facilitate better decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。