arXiv:2505.22777cs.CL2025-05Conference of the …被引 1

构建多语言对话评估框架,发现主流大模型评语缺陷。

MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Dialogue Evaluators

  • 用多模型生成多语言对话,动态构建评估数据
  • 发现顶级模型在共情、常识等维度评估失效
  • 适合评估大模型评测能力的研究者使用

当前开放域聊天机器人评价日益依赖大模型作为自动评判者。然而现有元评估基准存在静态、过时、多语言覆盖不足等问题,难以捕捉评估中的细微缺陷。我们提出MEDAL,一个自动化多智能体框架,用于构建更具代表性和多样性的开放域对话评估基准。该框架利用多个先进大模型生成基于不同种子情境的多语言用户-聊天机器人对话,再以强模型GPT-4.1进行多维度分析,揭示显著的跨语言性能差异。基于大规模评估结果,我们构建新的多语言元评估基准,并人工标注样本以获取细致的质量判断。该基准用于评估多个推理与非推理类大模型作为对话评估者的能力。通过MEDAL,我们发现当前最先进评判模型无法可靠检测共情缺失、常识错误或相关性不足等细微问题。

原文摘要 · Abstract (English)

Evaluating the quality of open-domain chatbots has become increasingly reliant on LLMs acting as automatic judges. However, existing meta-evaluation benchmarks are static, outdated, and lacking in multilingual coverage, limiting their ability to fully capture subtle weaknesses in evaluation. We introduce MEDAL, an automated multi-agent framework for curating more representative and diverse open-domain dialogue evaluation benchmarks. Our approach leverages several state-of-the-art LLMs to generate user-chatbot multilingual dialogues, conditioned on varied seed contexts. Then, a strong LLM (GPT-4.1) is used for a multidimensional analysis of the performance of the chatbots, uncovering noticeable cross-lingual performance differences. Guided by this large-scale evaluation, we curate a new meta-evaluation multilingual benchmark and human-annotate samples with nuanced quality judgments. This benchmark is then used to assess the ability of several reasoning and non-reasoning LLMs to act as evaluators of open-domain dialogues. Using MEDAL, we uncover that state-of-the-art judges fail to reliably detect nuanced issues such as lack of empathy, commonsense, or relevance.

大模型评估多语言对话系统元评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。