arXiv:2410.11324cs.AIcs.CV2024-10被引 5

用合成数据提升扩散模型在复杂推理任务中的决策能力

Diffusion-Based Offline RL for Improved Decision-Making in Augmented ARC Task

  • 构建规则化生成的合成数据集SOLAR,解决ARC任务数据不足问题
  • 在简单ARC任务上,扩散基离线强化学习实现多步决策并正确识别答案状态
  • 适合研究离线强化学习与智能推理结合的学者参考

有效长时策略使人工智能系统能在复杂环境中通过序列决策进行导航。强化学习(RL)代理通过优化序列决策以最大化奖励,即使没有即时反馈。为验证基于扩散的离线强化学习方法LDCQ在多步决策中展现的强推理能力,我们将其应用于抽象与推理基准测试集(Abstraction and Reasoning Corpus, ARC)。然而,由于ARC训练集中经验数据不足,将离线强化学习应用于提升AI的策略推理能力面临挑战。为此,我们提出了增强型离线强化学习数据集SOLAR(Synthesized Offline Learning Data for Abstraction and Reasoning),以及SOLAR-Generator,该生成器基于预定义规则生成多样化轨迹数据。SOLAR提供了充足的体验数据,使离线强化学习方法得以应用。我们在一个简单任务上合成SOLAR,并使用LDCQ方法训练智能体。实验表明,该方法在简单ARC任务上有效,智能体具备多步序列决策能力并能正确识别答案状态。结果凸显了离线强化学习在提升AI战略推理能力方面的潜力。

原文摘要 · Abstract (English)

Effective long-term strategies enable AI systems to navigate complex environments by making sequential decisions over extended horizons. Similarly, reinforcement learning (RL) agents optimize decisions across sequences to maximize rewards, even without immediate feedback. To verify that Latent Diffusion-Constrained Q-learning (LDCQ), a prominent diffusion-based offline RL method, demonstrates strong reasoning abilities in multi-step decision-making, we aimed to evaluate its performance on the Abstraction and Reasoning Corpus (ARC). However, applying offline RL methodologies to enhance strategic reasoning in AI for solving tasks in ARC is challenging due to the lack of sufficient experience data in the ARC training set. To address this limitation, we introduce an augmented offline RL dataset for ARC, called Synthesized Offline Learning Data for Abstraction and Reasoning (SOLAR), along with the SOLAR-Generator, which generates diverse trajectory data based on predefined rules. SOLAR enables the application of offline RL methods by offering sufficient experience data. We synthesized SOLAR for a simple task and used it to train an agent with the LDCQ method. Our experiments demonstrate the effectiveness of the offline RL approach on a simple ARC task, showing the agent's ability to make multi-step sequential decisions and correctly identify answer states. These results highlight the potential of the offline RL approach to enhance AI's strategic reasoning capabilities.

离线强化学习扩散模型推理任务数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。