用自然语言表达强化学习价值,让智能体更懂环境、主动思考。
Natural Language Reinforcement Learning
- 把价值函数转为可读的语义叙述,替代传统标量值。
- 在4个复杂任务中实现高效策略训练,表现优于传统方法。
- 适合需要解释性与自主决策的复杂智能体系统。
人工智能正迈向「经验时代」,要求智能体从持续、具身的交互中学习。传统强化学习(RL)将价值表示为标量,限制了智能体对环境的深层理解,并阻碍主动、审慎的学习。为此,我们提出自然语言强化学习(NLRL),将强化学习核心概念扩展至自然语言层面。核心是语言价值函数(LVF),将价值重构为可解释的语言叙事,阐明评估背后的逻辑。NLRL进一步将该思想延伸至策略、贝尔曼方程和策略迭代等关键组件。借助大语言模型(LLMs)的进展,可实现无需标注的环境交互,完成类似强化学习的策略与价值训练。在4个多步代理任务上的实验表明,NLRL在有效性、效率方面均表现优异,具备促进深层理解与主动学习策略的潜力。
原文摘要 · Abstract (English)
Artificial intelligence progresses towards the "Era of Experience," where agents are expected to learn from continuous, grounded interaction. We argue that traditional Reinforcement Learning (RL), which typically represents value as a scalar, can restrict agent's deep understanding of environments and hinders the active, deliberative learning crucial for navigating this new paradigm. To address the issue, we introduce Natural Language Reinforcement Learning (NLRL), a framework that extends RL principles into natural language counterparts. Central to NLRL is the Language Value Function (LVF), which redefines value as an interpretable linguistic narrative articulating the rationale behind an evaluation. NLRL further extends this concept to core RL components, including policy, the Bellman equation, and policy iteration. Leveraging recent advancements in Large Language Models (LLMs), NLRL can be practically implemented to achieve RL-like policy and value training through unsupervised environment interactions. Experiments over 4 multi-step agentic tasks demonstrate NLRL's effectiveness, efficiency, and its potential to foster deeper understanding and more active learning strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。