arXiv:2603.13309cs.LGcs.AI2026-03被引 2

解决自进化推理系统中问题多样性退化问题,提升长期学习能力。

Preventing Curriculum Collapse in Self-Evolving Reasoning Systems

  • 以问题为中心设计持续多样性信号,避免生成重复题目
  • 在7个数学基准上6项超越基线,最高提纯3.98分
  • 适合研究自进化系统、数学推理与模型可持续优化的学者

自进化推理框架使大模型通过迭代生成和求解问题实现自主能力提升,依赖可验证奖励。理想情况下,系统应探索多样问题空间并提出高价值新挑战。然而,尽管表面形式变化存在,近期研究发现此类系统在数轮迭代后仍会出现问题多样性崩溃。本文提出Prism方法,通过嵌入空间的语义划分建立持久多样性信号,促进跨轮次对低覆盖区域的均衡探索,并结合近发展区(ZPD)门控机制保持适切难度。在七个主流数学推理基准上,相较于五种自进化基线,Prism在六项任务中表现最佳,相较R-Zero在AMC上提升3.98个百分点,在Minerva Math上提升3.68个百分点。Prism还生成了语义多样且具挑战性的题目,构建出包含10万道题的Prism-Math数据集。结果表明,跨轮次语义覆盖是提升自进化推理能力的关键且未被充分挖掘的方向。代码、数据集与模型已开源。

原文摘要 · Abstract (English)

Self-evolving reasoning frameworks let LLMs improve their reasoning capabilities by iteratively generating and solving problems without external supervision, using verifiable rewards. Ideally, such systems are expected to explore a diverse problem space and propose new challenges of high learning value. While prior work has largely focused on solver-side optimisation and verification, recent evidence suggests that self-evolving systems can exhibit diversity collapse in posing new problems after just a few iterations, even when surface-level variation is preserved. We introduce Prism, a question-centric self-evolution method that directly tackles this collapse. Prism defines a persistent diversity signal over an embedding-induced semantic partition of mathematical problems and uses it to encourage balanced exploration of underrepresented regions across iterations. This coverage signal is combined with a Zone-of-Proximal-Development (ZPD) gate to preserve edge-of-solvability difficulty. Evaluated on seven widely used mathematical reasoning benchmarks against five self-evolving baselines, Prism achieves the highest accuracy on six out of seven tasks, achieving gains of +3.98 absolute points over R-Zero on AMC and +3.68 on Minerva Math. Prism also generates semantically diverse and challenging questions across iterations, resulting in the construction of the Prism-Math dataset comprising 100k mathematical questions. These results demonstrate that cross-iteration semantic coverage is a high-leverage and under-explored axis for building more capable self-evolving reasoners. We release the code, dataset, and models to facilitate further research.

自进化数学推理多样性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。