arXiv:2608.30955cs.AI2026-08

在有限交互中高效学习带条件与量化效果的动作模型

Learning Action Models with Conditional and Quantified Effects via Uncertainty-Guided Exploration

论文配图:Learning Action Models with Conditional and Quantified Effects via Uncertainty-Guided Exploration
图 1 · 摘自论文原文
  • 基于假设驱动的在线探索,主动选择最能减少不确定性的动作
  • 在6个基准任务中用更少数据解决更多任务,抗噪声能力强
  • 适合机器人规划等需精准动作模型的现实场景

精确的动作模型对有效规划至关重要。现有方法大多依赖简单动作表示,或在学习条件性与量化效果时计算不可行。本文提出在线假设驱动的条件动作模型学习(OHCAM),一种从有限环境交互中学习此类动作模型的在线方法。OHCAM维护一组动作模型假设的信念,并通过最大化竞争假设间的分歧来主动选择信息量大的动作以降低不确定性,同时对观测噪声具有鲁棒性。为实现可扩展性,OHCAM从少量简单假设开始,仅当当前假设与数据不一致时才扩展至更复杂的条件。六组基准规划领域的实验表明,即使存在观测噪声,OHCAM在样本效率上显著优于基线,能解决更多任务。我们在Kinova Gen3机器人上验证了两个任务,证明该方法具备实际应用价值。

原文摘要 · Abstract (English)

Accurate action models are critical for effective planning. Existing action-model learning methods largely assume simple action representations or become computationally intractable when learning conditional and quantified effects. We present Online Hypothesis-Driven Conditional Action Model Learning (OHCAM), an online approach for learning action models with conditional and quantified effects from limited interactions with the environment. OHCAM maintains a belief over hypothesized action models and actively selects informative actions to reduce uncertainty by maximizing disagreement among competing hypotheses, while being robust to noisy observations. To enable scalability, OHCAM begins with a small set of simple action model hypotheses and expands to more complex conditions only when the current hypotheses become inconsistent with the data. Experiments on six benchmark planning domains demonstrate that OHCAM is sample efficient in learning action models that solve substantially more tasks than baselines, even with observation noise. We validate OHCAM on two tasks using a Kinova Gen3 robot, demonstrating the real-world applicability of our approach.

动作建模在线学习机器人规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。