arXiv:2608.14851cs.AIcs.LG2026-08

用离线强化学习自动发现高质量国际象棋谜题,提升初学者学习效果。

Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning

论文配图:Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning
图 1 · 摘自论文原文
  • 基于用户解题历史数据,用离线强化学习评估谜题教学价值。
  • 对100-1000分初学者效果显著,尤其帮助学习停滞者进步。
  • 专家评级验证新发现谜题具有高教学价值,适合教育系统应用。

学习与技能掌握依赖大量刻意练习。在诸多学习场景中,制作高质量教学材料需深厚领域知识且耗时费力。教学材料需引导学生训练不同思维模式。在国际象棋中,谜题用于练习下一步计算和棋形识别。为初学者设计涵盖不同策略、合理前瞻步数的谜题集极具挑战性。主流平台如Chess.com和Lichess提供数百万自动生成谜题,但因缺乏人工设计,常被认为教学价值低。这些平台仅依赖启发式推荐。本文利用长达一年的用户解题历史数据(共15亿条记录),通过离线强化学习方法,学习谜题的教育价值并自动筛选更优谜题组合以支持学习者。实验表明,该方法对解题等级100–1000的初学者有显著影响,尤其改善学习停滞群体。我们还邀请专家对模型发现的谜题进行标注评分,结果证实其高教学价值。本研究展示了仅凭通用用户交互数据即可理解练习项教学潜力的可行性。

原文摘要 · Abstract (English)

Learning and skill mastery require extensive and deliberate practice. In many learning settings, producing high-quality pedagogical materials can require a high level of domain expertise and be very time-consuming. Pedagogical materials often need to train students to engage in different thinking patterns. In some domains, such as chess, puzzles are used to help students practice their skills in calculating the next moves and recognizing known patterns on a board. Giving students a practice set of puzzles to help them learn different modes of thinking is challenging because the teacher needs to carefully balance between different motifs and how many look-ahead steps a student needs to perform. Popular online platforms like Chess.com and Lichess offer players millions of puzzles. Unlike chess tactics puzzles procured by human experts, where chess beginners can learn valuable insights, these puzzles are automatically generated and often regarded as having low pedagogical value. These platforms also rely on a heuristic to recommend puzzles to users for practice. Using the user history data over an entire year, a total of 1.5 billion puzzle-solving histories, we learn the pedagogical value of a puzzle and how to automatically choose a set of puzzles to better support chess learners using insights from offline reinforcement learning. We show that using offline policy evaluation, our trained policy has significant impact on beginners with puzzle-solving Elo range of 100--1000, particularly for the group of beginners whose learning growth was stagnant. We also performed a qualitative analysis of the puzzles discovered by our model by collecting annotation ratings from expert chess players. The success of our pipeline shows promise for a future where we can understand the pedagogical values of practice items given general user interaction data.

强化学习教育技术国际象棋离线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。