arXiv:2510.24684cs.CL2025-10被引 70

让AI自己出题自己解,持续提升推理能力

SPICE: Self-Play In Corpus Environments Improves Reasoning

  • 模型分角色:出题者挖掘文档生成难题,解答者破解问题
  • 数学与通用推理能力分别提升8.9%和9.8%
  • 适合想实现持续自我进化AI的研究者

自适应系统需环境交互以持续进化。我们提出SPICE(Self-Play In Corpus Environments),一种强化学习框架,单个模型扮演双重角色:挑战者从大规模语料库中挖掘文档生成多样化推理任务,解题者则解决这些任务。通过对抗性动态,挑战者在解题者能力边界处自动构建难度递增的课程,而语料库的外部信号提供了丰富且近乎无限的反馈,支持持续改进。相比现有无语境自对弈方法,SPICE在多个模型家族上实现了数学推理(+8.9%)与通用推理(+9.8%)的稳定提升。分析表明,文档接地是关键机制,使系统能持续生成并达成更具挑战性的目标,从而实现长效自进化。

原文摘要 · Abstract (English)

Self-improving systems require environmental interaction for continuous adaptation. We introduce SPICE (Self-Play In Corpus Environments), a reinforcement learning framework where a single model acts in two roles: a Challenger that mines documents from a large corpus to generate diverse reasoning tasks, and a Reasoner that solves them. Through adversarial dynamics, the Challenger creates an automatic curriculum at the frontier of the Reasoner's capability, while corpus grounding provides the rich, near-inexhaustible external signal necessary for sustained improvement. Unlike existing ungrounded self-play methods that offer more limited benefits, SPICE achieves consistent gains across mathematical (+8.9%) and general reasoning (+9.8%) benchmarks on multiple model families. Our analysis reveals how document grounding is a key ingredient in SPICE to continuously generate its own increasingly challenging goals and achieve them, enabling sustained self-improvement.

自进化推理增强强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。