arXiv:2609.07575cs.LGcs.AI2026-09

只靠内在探索就能自发产生复杂行为,无需外部奖励。

Efficient Exploration Is Enough

论文配图:Efficient Exploration Is Enough
图 1 · 摘自论文原文
  • 以可泛化经验为优先,让智能体主动找最有学习价值的区域。
  • 理论证明最优探索者会先访问信息量最大的区域。
  • 实验显示即使简单环境也能自动生成渐进式行为课程。

本文提出一种高效探索的新视角,研究其在无外在奖励时的理论与实证影响。我们定义高效探索者为优先生成可泛化经验的智能体,即能支持跨环境预测与适应的学习数据。由此,我们从预测与泛化角度分析探索机制。理论上,最优探索者会自然规划轨迹,优先访问最具信息量和可学习性的区域。实证上,优化此类目标可催生自动化的渐进式行为课程,即便在相对简单的环境中亦然。结果表明,仅追求这一内在目标即可引发高度复杂行为的涌现。我们认为该框架为智能体-环境系统在无外部奖励、任务或目标下持续发展复杂行为提供了原则性机制。

原文摘要 · Abstract (English)

This work introduces an alternative view of efficient exploration and studies its theoretical and empirical implications in the absence of extrinsic rewards. Specifically, we define efficient explorers as agents that prioritize generating generalizable experience, i.e., data that supports learning models capable of predicting and adapting across the environment. This allows us to analyze efficient exploration through the lens of prediction and generalization. Theoretically, we demonstrate that optimally efficient explorers naturally schedule their trajectories to visit the most informative and learnable regions first. Empirically, we show that optimizing for these agents gives rise to an automatic curriculum of progressively more complex behaviors, even in relatively simple environments. These results indicate that pursuing this purely intrinsic objective alone is enough to drive the emergence of highly sophisticated behaviors. We believe that this new framework provides a principled mechanism by which agent-environment systems may sustain an open-ended process of increasingly complex behavior without external rewards, tasks, or objectives.

强化学习内在探索自适应行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。