arXiv:2609.07461cs.LG2026-09

让智能体在隐藏阶段转换中自动推断时间因果,实现最优决策。

Temporal-Causal Inference for Reinforcement Learning via Automata Learning

  • 用有限状态自动机建模隐藏阶段转换,通过反例驱动优化
  • 在基因治疗和交通信号场景中恢复正确因果模型并达最优性能
  • 适合处理不可逆阶段转移的非马尔可夫决策问题

我们研究存在不可逆阶段转换且由隐藏时间模式驱动的环境中的强化学习问题。智能体只能观测基础状态,无法直接感知当前阶段。我们将该问题形式化为两阶段非马尔可夫决策过程,并提出时序因果推断强化学习(TCIRL)框架,联合学习控制策略并推断阶段转换的隐藏时序原因。TCIRL维护一个假设的确定性有限自动机(DFA),用于追踪当前活跃阶段,并通过基于SAT的反例驱动合成不断修正。理论上证明,该假设几乎必然收敛至能识别真实因果语言的DFA,从而获得原非马尔可夫决策过程的最优策略。在基因治疗网格世界与交通信号环境中实验表明,TCIRL成功恢复正确的因果DFA,并达到全信息基准的性能水平。

原文摘要 · Abstract (English)

We consider reinforcement learning in environments with dynamics that undergo an irreversible phase transition governed by a hidden temporal pattern. The agent observes the base state but cannot observe the phase directly. We formalize this problem as a two-phase non-Markovian decision process and introduce Temporal-Causal Inference for Reinforcement Learning (TCIRL), a framework that jointly learns a control policy and infers the hidden temporal cause of the phase transition. TCIRL maintains a hypothesis deterministic finite automaton (DFA) to track what phase is active and refines it via counterexample-driven SAT-based synthesis. We prove that the hypothesis converges almost surely to a DFA recognizing the true cause language on all attainable label sequences, yielding an optimal policy for the original non-Markovian decision process. Experiments on a genetic therapy gridworld and a traffic signal environment show that TCIRL recovers the correct cause DFA and matches the full-information baseline in both domains.

强化学习时序推理因果推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。