arXiv:2501.15618cs.ROcs.AI2025-01中稿 · RLC 2025被引 6

逆约束学习实际学到的是不可逆失败状态的后向可达管,而非已失败区域。

Your Learned Constraint is Secretly a Backward Reachable Tube

  • 通过逆向推导安全示范,发现约束本质是后向可达管
  • 所学约束依赖数据采集系统的动态特性,影响策略搜索效率
  • 适合关注安全控制迁移与样本效率的研究者

逆约束学习(ICL)旨在从安全示范中推断约束条件,以期用于新任务的安全策略搜索,甚至在不同动力学下迁移。本文揭示了一个关键发现:无论理论上还是实践中,ICL 实际恢复的是失败不可避免的状态集合,而非已发生失败的状态集。在安全控制语境中,这对应于后向可达管(BRT),而非失败集。与失败集不同,BRT 依赖于数据收集系统的动力学特性。该动态依赖性对策略搜索的样本效率以及学习约束的可迁移性均产生重要影响。

原文摘要 · Abstract (English)

Inverse Constraint Learning (ICL) is the problem of inferring constraints from safe (i.e., constraint-satisfying) demonstrations. The hope is that these inferred constraints can then be used downstream to search for safe policies for new tasks and, potentially, under different dynamics. Our paper explores the question of what mathematical entity ICL recovers. Somewhat surprisingly, we show that both in theory and in practice, ICL recovers the set of states where failure is inevitable, rather than the set of states where failure has already happened. In the language of safe control, this means we recover a backwards reachable tube (BRT) rather than a failure set. In contrast to the failure set, the BRT depends on the dynamics of the data collection system. We discuss the implications of the dynamics-conditionedness of the recovered constraint on both the sample-efficiency of policy search and the transferability of learned constraints.

安全控制逆约束学习后向可达管策略迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。