arXiv:2501.13241cs.LG2025-01中稿 · Transactions on Ma…被引 2

用条件扩散模型提升决策中未见组合状态的零样本泛化能力

State Combinatorial Generalization In Decision Making With Conditional Diffusion Models

  • 用条件扩散模型模仿成功轨迹,学习状态组合的通用表示
  • 在迷宫、驾驶和多智能体环境中,新组合状态的泛化性能显著优于传统强化学习
  • 适用于需要处理复杂状态组合的自动驾驶等现实决策场景

许多现实决策问题具有组合性质,状态(如自动驾驶汽车周围的交通)可视为基本元素(如行人、树木、其他车辆)的组合。由于组合复杂性,训练集中无法覆盖所有元素组合,导致一个关键但未被充分研究的问题:对已见元素的新组合状态实现零样本泛化。本文首次形式化该问题,并证明现有基于价值的强化学习算法因在未见状态中价值预测不可靠而表现不佳。我们指出仅靠探索无法解决此问题,需更表达能力强且泛化性好的模型。实验表明,使用条件扩散模型对成功轨迹进行行为克隆,在新组合状态上的泛化性能优于传统RL方法。在迷宫、驾驶和多智能体环境中的测试验证了该方法的优越性和广泛适用性。

原文摘要 · Abstract (English)

Many real-world decision-making problems are combinatorial in nature, where states (e.g., surrounding traffic of a self-driving car) can be seen as a combination of basic elements (e.g., pedestrians, trees, and other cars). Due to combinatorial complexity, observing all combinations of basic elements in the training set is infeasible, which leads to an essential yet understudied problem of zero-shot generalization to states that are unseen combinations of previously seen elements. In this work, we first formalize this problem and then demonstrate how existing value-based reinforcement learning (RL) algorithms struggle due to unreliable value predictions in unseen states. We argue that this problem cannot be addressed with exploration alone, but requires more expressive and generalizable models. We demonstrate that behavior cloning with a conditioned diffusion model trained on successful trajectory generalizes better to states formed by new combinations of seen elements than traditional RL methods. Through experiments in maze, driving, and multiagent environments, we show that conditioned diffusion models outperform traditional RL techniques and highlight the broad applicability of our problem formulation.

决策生成扩散模型组合泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。