用数学框架证明语言反馈能大幅提升智能体学习效率。
Formalizing Learning from Language Feedback with Provable Guarantees
- 提出语言反馈学习的理论框架,定义信息复杂度度量
- 证明丰富语言反馈可使学习速度指数级快于仅靠奖励
- 设计无后悔算法Helix,适用于不可靠大模型提示场景
交互式学习从观察和语言反馈中获取知识是受大型语言模型代理兴起推动的热门方向。尽管已有显著的实证成果,但此类决策问题仍缺乏严谨的理论基础。本文形式化了语言反馈学习(LLF)问题,提出了足以在隐式奖励下实现学习的充分假设,并引入“迁移消除维数”作为刻画LLF难度的度量。我们形式化了语言反馈中的信息决定学习复杂度的直觉,并展示了在某些情况下,仅依赖语言反馈的学习速度可比仅依赖奖励的学习快一个指数级。我们设计了一种无后悔算法HELiX,通过序列交互可保证求解LLF问题,其性能保证与迁移消除维数成比例。在多个实验领域中,我们验证了HELiX即使在反复提示大语言模型不稳定的条件下依然表现良好。本工作为基于通用语言反馈设计有理论保障的交互式学习算法迈出了关键一步。
原文摘要 · Abstract (English)
Interactively learning from observation and language feedback is an increasingly studied area driven by the emergence of large language model (LLM) agents. Despite impressive empirical demonstrations, so far a principled framing of these decision problems remains lacking. We formalize the Learning from Language Feedback (LLF) problem, assert sufficient assumptions to enable learning despite latent rewards, and introduce $\textit{transfer eluder dimension}$ as a measure to characterize the hardness of LLF. We formalize the intuition that information in the language feedback governs the learning complexity, and demonstrate cases where learning from rich language feedback can be exponentially faster than learning from reward. We develop a no-regret algorithm, called $\texttt{HELiX}$, that provably solves LLF problems through sequential interactions, with performance guarantees that scale with the transfer eluder dimension. Across several empirical domains, we show that $\texttt{HELiX}$ performs well even when repeatedly prompting LLMs does not work reliably. Our contributions mark an important step towards designing principled interactive learning algorithms using generic language feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。