arXiv:2411.11504cs.AIcs.CL2024-11被引 13

用自动验证器为大模型提供反馈,提升后训练效果。

Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering

  • 设计搜索-验证-反馈三阶段框架,用自动化验证器改进大模型。
  • 通过验证反馈机制显著提升模型输出可靠性与能力。
  • 适合研究大模型后训练、智能验证系统的学者参考。

机器学习的发展越来越重视强大模型和可扩展的监督信号。然而,基础模型的出现带来了有效监督信号不足的挑战,亟需探索新的监督信号和技术方法。本文提出验证器工程(verifier engineering),一种专为大模型时代设计的新型后训练范式。其核心是利用一系列自动化验证器执行验证任务,并向基础模型提供有意义的反馈。我们系统地将验证器工程流程分为三个关键阶段:搜索、验证和反馈,并对各阶段的前沿研究进展进行全面综述。我们认为,验证器工程是迈向通用人工智能的重要路径。

原文摘要 · Abstract (English)

The evolution of machine learning has increasingly prioritized the development of powerful models and more scalable supervision signals. However, the emergence of foundation models presents significant challenges in providing effective supervision signals necessary for further enhancing their capabilities. Consequently, there is an urgent need to explore novel supervision signals and technical approaches. In this paper, we propose verifier engineering, a novel post-training paradigm specifically designed for the era of foundation models. The core of verifier engineering involves leveraging a suite of automated verifiers to perform verification tasks and deliver meaningful feedback to foundation models. We systematically categorize the verifier engineering process into three essential stages: search, verify, and feedback, and provide a comprehensive review of state-of-the-art research developments within each stage. We believe that verifier engineering constitutes a fundamental pathway toward achieving Artificial General Intelligence.

后训练验证器大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。