用合成反馈降低大模型评估的人工成本,同时保持结果无偏。
Accelerating Unbiased LLM Evaluation via Synthetic Feedback
- 融合真人与合成反馈,统计上保证评估无偏
- 最多减少24.8%人工标注量,且可预测节省效果
- 无需调参,适用于大规模模型评估
在开发新大型语言模型时,常需通过外部反馈计算其胜率以评估性能。人类反馈是衡量连贯性、可读性和对齐度等细微特质的金标准,但成本高昂,且影响用户体验。合成反馈(由其他大模型或奖励模型生成)可替代人工标注,但可能引入偏差。本文提出一种统计上严谨的框架,结合人类与合成反馈,在减少人工标注的同时保持胜率计算的无偏性。实验表明,使用现成合成评估器可减少12.2%的人工标注,微调版本更可达24.8%。该方法具备通用性、可扩展性,且无需超参数调优,标注节省量可基于数据特征预估。
原文摘要 · Abstract (English)
When developing new large language models (LLMs), a key step is evaluating their final performance, often by computing the win-rate against a reference model based on external feedback. Human feedback is the gold standard, particularly for capturing nuanced qualities like coherence, readability, and alignment with human expectations. However, human evaluations are costly -- even for large tech companies -- and when conducted with active users, they may negatively impact user experience. A promising alternative is synthetic feedback, where evaluations are conducted by other large language models, including reward models. While this eliminates the need for costly human annotations, it introduces biases that may distort the evaluation process. In this work, we propose a statistically principled framework that integrates human and synthetic feedback to reduce reliance on human annotations while maintaining unbiased win-rate calculations. Our experiments demonstrate a reduction in human annotations by up to 12.2% with an off-the-shelf synthetic evaluator and up to 24.8% with a finetuned variant. Apart from being generalizable, scalable, and free of hyper-parameter tuning, our method offers predictable annotation savings, which can be estimated based on data-dependent characteristics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。