arXiv:2411.04424cs.CLcs.AI2024-11EMNLP被引 16

用贝叶斯方法校准大模型评估的胜率,让自动文本质量评估更准确。

Bayesian Calibration of Win Rate Estimation with LLM Evaluators

  • 基于贝叶斯推断设计两种校准方法,减少评估偏差。
  • 在6个数据集上验证,显著提升胜率估计准确性。
  • 适合需要可靠自动评估的模型开发与对比研究者。

近期大语言模型(LLMs)的发展展示了将其作为评估器来衡量生成文本质量的潜力。然而,直接使用LLM评估器比较不同系统时,可能因评估器固有的胜率估计偏差导致结果不可靠。为此,我们提出两种校准方法:贝叶斯胜率采样(BWRS)和贝叶斯Dawid-Skene,均利用贝叶斯推断更准确地推断生成式语言模型的真实胜率。我们在涵盖故事生成、摘要和指令遵循任务的六个数据集上进行了实证验证,结果表明两种方法均能有效提升使用LLM作为评估器时的胜率估计精度,为可靠的自动文本质量评估提供了有前景的方向。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) show the potential of using LLMs as evaluators for assessing the quality of text generations from LLMs. However, applying LLM evaluators naively to compare or judge between different systems can lead to unreliable results due to the intrinsic win rate estimation bias of LLM evaluators. In order to mitigate this problem, we propose two calibration methods, Bayesian Win Rate Sampling (BWRS) and Bayesian Dawid-Skene, both of which leverage Bayesian inference to more accurately infer the true win rate of generative language models. We empirically validate our methods on six datasets covering story generation, summarization, and instruction following tasks. We show that both our methods are effective in improving the accuracy of win rate estimation using LLMs as evaluators, offering a promising direction for reliable automatic text quality evaluation.

大模型评估贝叶斯推断胜率校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。