用分层语义重定义信号时序逻辑,揭示强化学习嵌入空间的几何结构。
Stratifying Reinforcement Learning with Signal Temporal Logic
- 将原子命题视为分层空间中的成员判断,使STL公式可诱导时空分层
- 在迷你网格游戏中验证了嵌入空间存在可识别的分层结构
- 为分析深度强化学习的表示提供新视角,适合研究表示几何的学者
本文提出一种基于分层的信号时序逻辑(STL)语义,其中每个原子命题被解释为在分层空间中的成员测试。这一视角揭示了分层理论与STL之间新颖的对应关系,表明大多数STL公式可视为诱导时空分层。该解释具有双重意义:首先,为分析深度强化学习(DRL)生成的嵌入空间结构提供了新的理论框架,并将其与环境决策空间的几何特性相联系;其次,提供了一个严谨的框架,既支持现有高维分析工具的复用,也推动新型计算技术的开发。为验证理论,我们(1)展示了分层理论在迷你网格游戏中的作用,(2)对一个在该游戏中运行的DRL智能体的隐空间嵌入应用数值技术,以STL公式的鲁棒性作为奖励。在此过程中,我们提出了计算高效的特征签名,初步证据显示其在揭示此类嵌入空间分层结构方面颇具潜力。
原文摘要 · Abstract (English)
In this paper, we develop a stratification-based semantics for Signal Temporal Logic (STL) in which each atomic predicate is interpreted as a membership test in a stratified space. This perspective reveals a novel correspondence principle between stratification theory and STL, showing that most STL formulas can be viewed as inducing a stratification of space-time. The significance of this interpretation is twofold. First, it offers a fresh theoretical framework for analyzing the structure of the embedding space generated by deep reinforcement learning (DRL) and relates it to the geometry of the ambient decision space. Second, it provides a principled framework that both enables the reuse of existing high-dimensional analysis tools and motivates the creation of novel computational techniques. To ground the theory, we (1) illustrate the role of stratification theory in Minigrid games and (2) apply numerical techniques to the latent embeddings of a DRL agent playing such a game where the robustness of STL formulas is used as the reward. In the process, we propose computationally efficient signatures that, based on preliminary evidence, appear promising for uncovering the stratification structure of such embedding spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。