让智能体在不可逆环境中安全学习,通过求助导师实现高效成长。
Safe Learning Under Irreversible Dynamics via Asking for Help
- 引入导师求助机制,避免不可逆错误。
- 在任意马尔可夫决策过程上,后悔和求助次数均亚线性增长。
- 适合高风险场景下的自主智能体,如医疗、航天等。
大多数具有正式后悔保证的学习算法依赖于尝试所有可能行为,但在某些错误无法恢复的场景下存在风险。本文提出允许学习智能体向导师求助,并在相似状态间转移知识。该组合使智能体既能安全又能高效学习。在标准在线学习假设下,我们设计了一种算法,在任意马尔可夫决策过程(MDP)中,包括具有不可逆动态的MDP,其后悔和导师查询次数均关于时间窗口为亚线性。证明过程包含一系列三重归约,可能具有独立研究价值。概念上,这是首个形式化证明:智能体可在未知、无界且高风险环境中获得高回报,并逐步实现自立,而无需系统重置。
原文摘要 · Abstract (English)
Most learning algorithms with formal regret guarantees essentially rely on trying all possible behaviors, which is problematic when some errors cannot be recovered from. Instead, we allow the learning agent to ask for help from a mentor and to transfer knowledge between similar states. We show that this combination enables the agent to learn both safely and effectively. Under standard online learning assumptions, we provide an algorithm whose regret and number of mentor queries are both sublinear in the time horizon for any Markov Decision Process (MDP), including MDPs with irreversible dynamics. Our proof involves a sequence of three reductions which may be of independent interest. Conceptually, our result may be the first formal proof that it is possible for an agent to obtain high reward while becoming self-sufficient in an unknown, unbounded, and high-stakes environment without resets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。