arXiv:2510.17928cs.LGcs.AI2025-10被引 1

用进化策略自动生成可验证的训练数据,提升模型泛化能力

EvoSyn: Generalizable Evolutionary Data Synthesis for Verifiable Learning

  • 通过进化框架联合生成问题、解法和验证机制
  • 在代码和代理任务上性能显著优于基线
  • 无需领域规则,适合数学、编程等多任务训练

可靠的可验证数据已成为现代语言模型能力提升的关键驱动力,支持稳定的强化学习与有效的知识迁移。然而,构建通用的合成可验证数据仍面临挑战:生成易幻觉,验证机制弱或平凡,难以区分优劣解。现有方法依赖特定任务启发式或事后过滤,缺乏普适的可验证性评估器。本文提出一种无任务依赖、策略引导、可执行验证的数据合成框架,仅需少量种子监督,即可协同生成问题、多样解法与验证机制,并通过一致性评估器迭代发现有效策略。该流程将过滤升级为原则性合成,可靠构建连贯且可验证的训练样本,实现跨域泛化。实验表明,在RLVR与模型蒸馏范式下,使用本框架生成的数据显著提升LiveCodeBench与AgentBench-OS任务表现,验证了其强大泛化能力。

原文摘要 · Abstract (English)

Reliable verifiable data has become a key driver of capability gains in modern language models, enabling stable reinforcement learning with verifiable rewards and effective distillation that transfers competence across math, coding, and agentic tasks. Yet constructing generalizable synthetic verifiable data remains difficult due to hallucination-prone generation, and weak or trivial verification artifacts that fail to separate strong from weak solutions. Existing approaches often rely on task-specific heuristics or post-hoc filters that do not transfer across domains and lack a principled, universal evaluator of verifiability. In this work, we introduce an evolutionary, task-agnostic, strategy-guided, executably-checkable data synthesis framework that, from minimal seed supervision, jointly synthesizes problems, diverse candidate solutions, and verification artifacts, and iteratively discovers strategies via a consistency-based evaluator that enforces agreement between human-annotated and strategy-induced checks. This pipeline upgrades filtering into principled synthesis: it reliably assembles coherent, verifiable training instances and generalizes without domain-specific rules. Our experiments demonstrate the effectiveness of the proposed approach under both RLVR and model distillation training paradigms. The results show that training with our synthesized data yields significant improvements on both the LiveCodeBench and AgentBench-OS tasks, highlighting the robust generalization of our framework.

数据合成可验证性强化学习泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。