用强化学习玩韩语接龙,发现奖励冲突并用课程学习缓解。
Studying the Korean Word-Chain Game with RLVR: Mitigating Reward Conflicts via Curriculum Learning
- 基于可验证奖励的强化学习训练模型玩韩语接龙。
- 规则奖励存在冲突,影响模型学习效果。
- 采用课程学习策略有效缓解奖励冲突,适合多语言谜题研究。
基于可验证奖励的强化学习(RLVR)是一种有前景的方法,可用于训练具备更强推理能力的大语言模型(LLMs),并已应用于多种逻辑谜题。本文利用RLVR研究韩语接龙游戏,发现由规则生成的奖励存在自然冲突,并通过实验表明课程学习方案能有效缓解这些冲突。研究结果为在多种语言中开展谜题任务的研究提供了新思路。
原文摘要 · Abstract (English)
Reinforcement learning with verifiable rewards (RLVR) is a promising approach for training large language models (LLMs) with stronger reasoning abilities. It has also been applied to a variety of logic puzzles. In this work, we study the Korean word-chain game using RLVR. We show that rule-derived rewards can naturally conflict, and demonstrate through experiments that a curriculum-learning scheme mitigates these conflicts. Our findings motivate further studies of puzzle tasks in diverse languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。