arXiv:2508.09624cs.LGcs.AI2025-08

用因果能力发现关键目标,让智能体更高效探索环境。

Goal Discovery with Causal Capacity for Efficient Reinforcement Learning

  • 提出因果容量度量状态空间中行为对轨迹的影响
  • 高因果容量状态对应预期子目标,提升探索成功率
  • 适合需要高效探索的复杂强化学习任务

因果推理对人类探索世界至关重要,可建模为智能体在强化学习中高效探索环境的机制。现有研究指出,建立动作与状态转移间的因果关系,有助于智能体推理策略对未来轨迹的影响,从而实现有目的的探索。然而,在复杂场景的庞大状态-动作空间中,因果性难以衡量。本文提出一种新的目标发现框架——因果容量引导的目标发现(GDCC),首先推导出状态空间中的因果容量度量,表示智能体行为对未来轨迹的最大影响。随后,提出基于蒙特卡洛的方法识别离散状态空间中的关键点,并进一步优化以适应连续高维环境。这些关键点揭示了智能体在环境中做出重要决策的位置,被作为子目标,引导智能体更目的性、高效地进行探索。多目标任务的实验结果表明,高因果容量状态与预期子目标高度一致,且GDCC相比基线方法显著提升了成功率达数个百分点。

原文摘要 · Abstract (English)

Causal inference is crucial for humans to explore the world, which can be modeled to enable an agent to efficiently explore the environment in reinforcement learning. Existing research indicates that establishing the causality between action and state transition will enhance an agent to reason how a policy affects its future trajectory, thereby promoting directed exploration. However, it is challenging to measure the causality due to its intractability in the vast state-action space of complex scenarios. In this paper, we propose a novel Goal Discovery with Causal Capacity (GDCC) framework for efficient environment exploration. Specifically, we first derive a measurement of causality in state space, \emph{i.e.,} causal capacity, which represents the highest influence of an agent's behavior on future trajectories. After that, we present a Monte Carlo based method to identify critical points in discrete state space and further optimize this method for continuous high-dimensional environments. Those critical points are used to uncover where the agent makes important decisions in the environment, which are then regarded as our subgoals to guide the agent to make exploration more purposefully and efficiently. Empirical results from multi-objective tasks demonstrate that states with high causal capacity align with our expected subgoals, and our GDCC achieves significant success rate improvements compared to baselines.

强化学习因果推理目标发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。