用因果图时序逻辑加速长期任务的强化学习。
Inferring Causal Graph Temporal Logic Formulas to Expedite Reinforcement Learning in Temporally Extended Tasks
- 联合学习策略与因果图时序逻辑规范,构建闭环优化框架。
- 在基因与电力网络中实现更快收敛,行为更可验证。
- 适合需要高效探索与可解释决策的复杂动态系统研究者。
决策任务常表现为具有时空动态的图结构。黑箱强化学习往往忽视局部变化在网络中的传播机制,导致样本效率低且可解释性差。本文提出GTL-CIRL闭环框架,同时学习策略并挖掘因果图时序逻辑(Causal GTL)规范。该方法通过增强鲁棒性的奖励设计,当效果失败时收集反例,并利用高斯过程驱动的贝叶斯优化来精炼参数化因果模板。高斯过程模型捕捉系统动态中的空间与时间相关性,从而高效探索复杂参数空间。在基因网络与电力网络的案例研究中,相比标准强化学习基线,该方法实现更快的学习速度和更清晰、可验证的行为表现。
原文摘要 · Abstract (English)
Decision-making tasks often unfold on graphs with spatial-temporal dynamics. Black-box reinforcement learning often overlooks how local changes spread through network structure, limiting sample efficiency and interpretability. We present GTL-CIRL, a closed-loop framework that simultaneously learns policies and mines Causal Graph Temporal Logic (Causal GTL) specifications. The method shapes rewards with robustness, collects counterexamples when effects fail, and uses Gaussian Process (GP) driven Bayesian optimization to refine parameterized cause templates. The GP models capture spatial and temporal correlations in the system dynamics, enabling efficient exploration of complex parameter spaces. Case studies in gene and power networks show faster learning and clearer, verifiable behavior compared to standard RL baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。