通过几何共识机制让模型自演化,提升科学推理的可靠性与多样性。
Sci-CoE: Co-evolving Scientific Reasoning LLMs via Geometric Consensus with Sparse Supervision
- 两阶段自演化框架:先用少量标注数据建立验证基准,再在无监督下迭代优化。
- 几何奖励机制融合共识、可靠性和多样性,推动大规模模型自我迭代。
- 在多个科学推理基准上表现优异,适合构建更稳健的智能评估系统。
大型语言模型在推理任务中展现出卓越能力,协同进化范式在代码和数学领域已取得显著成果。然而,在科学推理任务中,这些模型仍因不可靠的解题评估和验证策略多样性不足而表现脆弱。本文提出Sci-CoE,一种两阶段科学协同进化框架,使模型在解题与验证角色间自我演化,从稀疏监督逐步过渡到无监督学习。第一阶段利用少量标注数据建立验证器的基本正确性判断基准;第二阶段引入几何奖励机制,综合考虑一致性、可靠性和多样性,驱动在无标签数据上的大规模自我迭代。在多个通用科学推理基准上的实验表明,Sci-CoE显著增强了复杂推理能力,具备良好可扩展性,有助于构建更鲁棒且多样化的评估体系。代码已开源:https://github.com/InternScience/Sci-CoE。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated exceptional reasoning capabilities, and co-evolving paradigms have shown promising results in domains such as code and math. However, in scientific reasoning tasks, these models remain fragile due to unreliable solution evaluation and limited diversity in verification strategies. In this work, we propose Sci-CoE, a two-stage scientific co-evolving framework that enables models to self-evolve as both solver and verifier through a transition from sparse supervision to unsupervised learning. In the first stage, the model uses a small set of annotated data to establish fundamental correctness judgment anchors for the Verifier. In the second stage, we introduce a geometric reward mechanism that jointly considers consensus, reliability, and diversity, driving large-scale self-iteration on unlabeled data. Experiments on several general scientific benchmarks demonstrate that Sci-CoE enhances complex reasoning capabilities and exhibits strong scalability, facilitating the construction of more robust and diverse evaluation systems. Codes are available at https://github.com/InternScience/Sci-CoE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。