arXiv:2505.09114cs.AIcs.LG2025-05IJCAI被引 2

用反事实推理让决策Transformer在数据少时仍能做出好决策

Beyond the Known: Decision Making with Counterfactual Reasoning Decision Transformer

  • 通过生成反事实经验,让模型在没学过的场景中也能推理
  • 在有限数据和动态变化环境下,性能超越传统决策Transformer
  • 无需改架构就能组合低质量轨迹,适合真实世界应用

决策Transformer(DT)在现代强化学习中扮演关键角色,利用离线数据在多个领域取得优异表现。然而,DT需要高质量、全面的数据才能发挥最佳性能。在现实应用中,训练数据稀缺且最优行为罕见,导致基于离线数据的训练充满挑战,次优数据会抑制性能。为此,我们提出反事实推理决策Transformer(CRDT),一种受反事实推理启发的新框架。CRDT通过生成和利用反事实经验,增强DT在已知数据之外的推理能力,从而在未见过的场景中实现更优决策。在Atari和D4RL基准上的实验表明,即使在数据有限和动态变化的场景下,CRDT也优于传统DT方法。此外,反事实推理使DT代理具备拼接能力,可组合次优轨迹而无需修改架构。这些结果凸显了反事实推理在提升强化学习智能体性能与泛化能力方面的潜力。

原文摘要 · Abstract (English)

Decision Transformers (DT) play a crucial role in modern reinforcement learning, leveraging offline datasets to achieve impressive results across various domains. However, DT requires high-quality, comprehensive data to perform optimally. In real-world applications, the lack of training data and the scarcity of optimal behaviours make training on offline datasets challenging, as suboptimal data can hinder performance. To address this, we propose the Counterfactual Reasoning Decision Transformer (CRDT), a novel framework inspired by counterfactual reasoning. CRDT enhances DT ability to reason beyond known data by generating and utilizing counterfactual experiences, enabling improved decision-making in unseen scenarios. Experiments across Atari and D4RL benchmarks, including scenarios with limited data and altered dynamics, demonstrate that CRDT outperforms conventional DT approaches. Additionally, reasoning counterfactually allows the DT agent to obtain stitching abilities, combining suboptimal trajectories, without architectural modifications. These results highlight the potential of counterfactual reasoning to enhance reinforcement learning agents' performance and generalization capabilities.

决策Transformer反事实推理强化学习离线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。