arXiv:2606.11127cs.CLcs.AI2026-06被引 1

提升合成数据筛选的可靠性与回收效率,让错误生成可追溯、可修复。

Provenance-Grounded Gating and Adaptive Recovery in Synthetic Post-Training Data Curation

论文配图:Provenance-Grounded Gating and Adaptive Recovery in Synthetic Post-Training Data Curation
图 1 · 摘自论文原文
  • 基于生成来源的溯源信息优化筛选机制,增强判断准确性。
  • 两类过滤器拒收样本重叠少,需协同使用以减少遗漏。
  • 动态诊断并针对性重生成失败样本,显著提高数据回收率。

合成后训练流程通常使用奖励模型或大语言模型判别器过滤生成内容,但两个关键问题长期未被共同考察:过滤信号是否基于生成背后的原始证据,以及被拒绝样本能否系统性恢复而非直接丢弃。本文通过对抗性注入语料库,在不同门控配置、恢复策略和生成器规模下开展受控实验,提供真实错误标签。结果表明,精确的源溯源信息能提升强判别器对忠实性的筛选效果;幻觉检测与奖励信号拒收群体几乎不重叠,二者缺一不可;结合故障诊断与定向重生成的自适应恢复流程,相比简单重采样,显著提升了数据产出量、回收率和注入样本召回率。下游微调质量主要由生成器规模决定,而过滤与恢复策略虽次之,仍具显著影响。

原文摘要 · Abstract (English)

Synthetic post-training pipelines commonly filter generated samples with reward models or holistic LLM judges, yet two practices remain rarely examined together: whether the filtering signal is grounded in the source evidence that induced each generation, and whether rejected samples can be systematically recovered rather than permanently discarded. We present a controlled study of both questions across gate configurations, recovery strategies, and generator scales, using adversarially injected corpora to provide ground-truth failure labels. We find that exact source provenance improves faithfulness gating for stronger judges, that hallucination and reward gates reject largely disjoint sample populations making both necessary, and that an adaptive recovery pipeline combining failure diagnosis with targeted regeneration achieves higher yield, recovery rate, and injection recall than naive resampling. Downstream fine-tuning quality is driven primarily by generator scale, with filtration and recovery conditions contributing meaningfully but secondarily.

合成数据数据筛选自适应恢复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。