arXiv:2505.19337cs.LGcs.AI2025-05

无需重训练,一键指定避让区域实现零样本避障决策

Prompting Decision Transformers for Zero-Shot Reach-Avoid Policies

  • 将目标与避让区直接作为提示词输入,动态生成策略
  • 仅用随机策略轨迹即实现零样本避障,最优基线提升35.7%成本性能
  • 适用于生物细胞重编程等复杂动态系统中的安全路径规划

离线目标条件强化学习在需抵达目标状态同时避开不良状态区域的避障任务中表现良好。现有方法通常将避让信息编码于状态空间和代价函数中,限制了评估时对新避让区域的灵活设定,且高度依赖精心设计的奖励与代价函数,难以扩展至复杂或结构不明确环境。本文提出RADT,一种离线、免奖励、目标与避让区域条件化的决策变压器模型。RADT通过提示词直接编码目标与避让区域,支持任意数量、任意大小的避让区域在评估时动态指定。仅利用随机策略产生的次优离线轨迹,通过创新的目标与避让区域回溯重标注机制,学习避障行为。我们在11个任务、环境及实验设置下对比3种现有离线目标条件强化学习模型,结果表明RADT可零样本泛化至分布外的避让区域大小与数量,超越需重新训练的基线。在一项零样本设置中,其标准化代价较最优重训练基线降低35.7%,同时保持高目标达成率。我们将RADT应用于生物学细胞重编程,有效减少轨迹中进入不良中间基因表达状态的次数,即使面对随机转移与离散结构化状态动力学。

原文摘要 · Abstract (English)

Offline goal-conditioned reinforcement learning methods have shown promise for reach-avoid tasks, where an agent must reach a target state while avoiding undesirable regions of the state space. Existing approaches typically encode avoid-region information into an augmented state space and cost function, which prevents flexible, dynamic specification of novel avoid-region information at evaluation time. They also rely heavily on well-designed reward and cost functions, limiting scalability to complex or poorly structured environments. We introduce RADT, a decision transformer model for offline, reward-free, goal-conditioned, avoid region-conditioned RL. RADT encodes goals and avoid regions directly as prompt tokens, allowing any number of avoid regions of arbitrary size to be specified at evaluation time. Using only suboptimal offline trajectories from a random policy, RADT learns reach-avoid behavior through a novel combination of goal and avoid-region hindsight relabeling. We benchmark RADT against 3 existing offline goal-conditioned RL models across 11 tasks, environments, and experimental settings. RADT generalizes in a zero-shot manner to out-of-distribution avoid region sizes and counts, outperforming baselines that require retraining. In one such zero-shot setting, RADT achieves 35.7% improvement in normalized cost over the best retrained baseline while maintaining high goal-reaching success. We apply RADT to cell reprogramming in biology, where it reduces visits to undesirable intermediate gene expression states during trajectories to desired target states, despite stochastic transitions and discrete, structured state dynamics.

决策变换器避障控制零样本学习生物建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。