arXiv:2508.00454cs.CL2025-08被引 1

用单模型模拟多专家评分,高效评估对话质量。

Learning an Efficient Multi-Turn Dialogue Evaluator from Multiple LLM Judges

  • 将多个大模型的评分知识融合到一个模型中
  • 在7个评测集上超越现有方法,速度更快
  • 适合需要快速、精准对话评估的研究者

评估大语言模型的对话能力仍具挑战。当前主流方法依赖“大模型作为评判者”范式,但易受各种偏差影响,降低评估可靠性。近期方法采用多个大模型作为评判者并聚合其意见,虽有效却带来巨大推理开销。本文提出一种高效对话评估器,通过整合多个大模型评判者的偏好知识,构建单一模型。该方法保留多评判者反馈的多样性优势,同时大幅降低评估成本,实现快速、灵活、细粒度的对话质量评估。在七个单评分与成对比较对话评估基准上的实验表明,该方法在多种场景下均优于现有基线,展现出高效性与鲁棒性。

原文摘要 · Abstract (English)

Evaluating the conversational abilities of large language models (LLMs) remains a challenging task. Current mainstream approaches primarily rely on the "LLM-as-a-judge" paradigm, where an LLM is prompted to serve as an evaluator to assess dialogue quality. However, such methods often suffer from various biases, which undermine the reliability and consistency of the evaluation results. To mitigate these biases, recent methods employ multiple LLMs as judges and aggregate their judgments to select the optimal assessment. Although effective, this multi-judge approach incurs significant computational overhead during inference. In this paper, we propose an efficient dialogue evaluator that captures the collective wisdom of multiple LLM judges by aggregating their preference knowledge into a single model. Our approach preserves the advantages of diverse multi-judge feedback while drastically reducing the evaluation cost, enabling fast, flexible, and fine-grained dialogue quality assessment. Extensive experiments on seven single rating and pairwise comparison dialogue evaluation benchmarks demonstrate that our method outperforms existing baselines across diverse scenarios, showcasing its efficiency and robustness.

对话评估多模型聚合效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。