研究目标导向行为对慢特征分析的影响及修正方法。
Slow Feature Analysis on Markov Chains from Goal-Directed Behavior
- 从马尔可夫链视角分析目标导向行为下的慢特征学习机制。
- 发现近奖励状态与远奖励状态的访问频率差异会损害价值函数逼近。
- 提出三种修正方案,适合关注强化学习表征设计的研究者。
慢特征分析是一种无监督表示学习方法,可从时序数据中提取变化缓慢的特征,作为后续强化学习的基础。通常假设生成数据的行为是均匀随机游走,但较少研究基于目标导向行为(强化学习中常见)的数据进行表征学习。在空间环境中,目标导向行为会导致靠近奖励位置与远离奖励位置的状态访问频率显著不同。本文从遍历马尔可夫链的最优慢特征视角,研究这种差异对价值函数逼近的负面影响,并评估和讨论三种可能缓解不利尺度效应的修正路径。此外,还考察了目标回避行为的特殊情况。
原文摘要 · Abstract (English)
Slow Feature Analysis is a unsupervised representation learning method that extracts slowly varying features from temporal data and can be used as a basis for subsequent reinforcement learning. Often, the behavior that generates the data on which the representation is learned is assumed to be a uniform random walk. Less research has focused on using samples generated by goal-directed behavior, as commonly the case in a reinforcement learning setting, to learn a representation. In a spatial setting, goal-directed behavior typically leads to significant differences in state occupancy between states that are close to a reward location and far from a reward location. Through the perspective of optimal slow features on ergodic Markov chains, this work investigates the effects of these differences on value-function approximation in an idealized setting. Furthermore, three correction routes, which can potentially alleviate detrimental scaling effects, are evaluated and discussed. In addition, the special case of goal-averse behavior is considered.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。