arXiv:2507.06582cs.LGcs.AI2025-07被引 1

用信息增益指导探索,高效学习环境的可控动态

Learning controllable dynamics through informative exploration

  • 基于预测信息增益选择最有价值的探索区域
  • 强化学习策略实现次优但可靠的探索,提升动态建模精度
  • 适合需要自主探索与建模的机器人和强化学习场景

具有可控动态的环境通常依赖显式模型来理解。然而,这类模型并不总能获得,有时可通过探索环境来学习。本文研究使用一种名为“预测信息增益”的信息度量,以确定下一步最具信息性的探索区域。结合强化学习方法,可找到表现良好的次优探索策略,从而获得对底层可控动态的可靠估计。该方法通过与多种短视探索策略对比得到验证。

原文摘要 · Abstract (English)

Environments with controllable dynamics are usually understood in terms of explicit models. However, such models are not always available, but may sometimes be learned by exploring an environment. In this work, we investigate using an information measure called "predicted information gain" to determine the most informative regions of an environment to explore next. Applying methods from reinforcement learning allows good suboptimal exploring policies to be found, and leads to reliable estimates of the underlying controllable dynamics. This approach is demonstrated by comparing with several myopic exploration approaches.

强化学习探索策略信息增益

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。