无需真实标签即可提升模型输出验证质量
FUSE: Ensembling Verifiers with Zero Labeled Data

- 通过控制验证器间依赖关系,实现无监督集成
- 在多个基准上性能媲美甚至超越有标签方法
- 适合缺乏标注数据的模型验证场景
模型输出验证正成为大语言模型训练与实际部署的关键环节。实践中,由于获取真实标签耗时费力,常使用不完善的LLM裁判和奖励模型。本文提出完全无监督评分集成方法FUSE,通过调控验证器间的条件依赖关系,提升一类谱系集成算法在无监督情形下的表现。该方法无需任何真实正确性标签,在多样化的生成模型、验证器和评测集上进行测试时,通常达到或超过半监督方法的性能。我们在传统学术基准如GPQA Diamond,以及前沿未饱和基准如Humanity's Last Exam和IMO Shortlist题目上验证了该方法的有效性。
原文摘要 · Abstract (English)
Verification of model outputs is rapidly emerging as a key primitive for both training and real-world deployment of large language models (LLMs). In practice, this often involves using imperfect LLM judges and reward models since ground truth acquisition can be time-consuming and expensive. We introduce Fully Unsupervised Score Ensembling (FUSE), a method for improving verification quality by ensembling verifiers without access to ground truth correctness labels. The key idea behind FUSE is to control conditional dependencies between verifiers in a manner that improves the unsupervised performance of a class of spectral algorithms from the ensembling literature. Despite requiring zero ground truth labels, FUSE typically matches or improves upon semi-supervised alternatives in test-time scaling experiments with diverse sets of generator models, verifiers, and benchmarks. In particular, we validate our method on both conventional academic benchmarks such as GPQA Diamond and on frontier, unsaturated benchmarks such as Humanity's Last Exam and IMO Shortlist questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。