从不完整信息中学习动作模型,突破了传统方法对状态完全可观的依赖。
Learning Lifted Action Models from Traces with Minimal Information About Actions and States
- 在仅知部分动作与状态信息下,设计三类可学习的算法框架
- 实验验证了在部分可观测条件下仍能还原等效动作域
- 适合研究自动规划与弱监督学习的学者参考
近期研究表明,仅凭动作轨迹即可高效准确地学习到提升版STRIPS模型(即无须观测状态);但该方法受限于动作参数冗余。后续工作通过引入隐式参数的STRIPS+模型缓解此问题,却要求状态完全可观测。本文进一步放宽条件,考虑更一般的场景:动作和状态均仅提供部分信息。我们提出了三类通用情形下的学习算法与完备性结论,分别假设仅部分动作参数可观、全可观部分状态谓词、局部可观部分状态谓词。对于给定的STRIPS+领域,这些结果刻画了从轨迹中学习等效领域的条件。实验结果支持了理论分析。
原文摘要 · Abstract (English)
It has been recently shown that lifted STRIPS models can be learned correctly and efficiently from action traces alone; i.e., applicable action sequences from a hidden STRIPS model. The result is remarkable because the states are not assumed to be observable at all, and yet it is not practical enough as STRIPS actions include arguments that are not needed for selecting the actions. This shortcoming has been addressed by assuming that the action traces come instead from a hidden STRIPS+ model where some action arguments are implicit in the hidden action preconditions. A limitation of this approach, however, is that it assumes that the states are fully observable. In this work, we relax these restrictions and consider the problem of learning STRIPS+ action domains from traces in a more general context where the traces carry partial information about both actions and states. In particular, we formulate algorithms and completeness results for three general cases, all of which assume full observability of selected action arguments. In the first case, no observability of the state is assumed; in the second case, full observability of some state predicates is assumed, and in the third case, local observability of some state predicates is assumed instead. Given a STRIPS+ domain, these results characterize the conditions under which an equivalent domain can be learned from traces. Experimental results are reported.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。