arXiv:2501.17399cs.CLcs.AI2025-01ACL被引 172

新基准挑战大模型多轮对话能力,顶尖模型准确率不足一半。

MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

  • 设计四类真实多轮对话难题,考验指令遵循与上下文推理。
  • 顶级模型如Claude 3.5 Sonnet平均准确率仅41.4%。
  • 自建语言模型评分机制,评估结果接近人工专家水平。

我们提出MultiChallenge,首个评估大语言模型在与人类用户进行多轮对话方面能力的基准,这一能力对实际应用至关重要但长期被忽视。MultiChallenge识别出四类常见且真实的多轮对话挑战,这些挑战不仅普遍存在于当前人机交互中,也对所有前沿大模型构成严峻考验。这四类挑战均需同时具备精准的指令遵循、上下文分配和上下文推理能力。我们还构建了基于大模型的评分系统,采用逐实例评分标准,实现与经验丰富的真人评审者高度一致的自动评估。尽管在现有评测基准上表现接近完美,所有前沿模型在MultiChallenge上的准确率均低于50%,其中表现最佳的Claude 3.5 Sonnet(2024年6月)平均准确率为41.4%。

原文摘要 · Abstract (English)

We present MultiChallenge, a pioneering benchmark evaluating large language models (LLMs) on conducting multi-turn conversations with human users, a crucial yet underexamined capability for their applications. MultiChallenge identifies four categories of challenges in multi-turn conversations that are not only common and realistic among current human-LLM interactions, but are also challenging to all current frontier LLMs. All 4 challenges require accurate instruction-following, context allocation, and in-context reasoning at the same time. We also develop LLM as judge with instance-level rubrics to facilitate an automatic evaluation method with fair agreement with experienced human raters. Despite achieving near-perfect scores on existing multi-turn evaluation benchmarks, all frontier models have less than 50% accuracy on MultiChallenge, with the top-performing Claude 3.5 Sonnet (June 2024) achieving just a 41.4% average accuracy.

多轮对话评估基准大模型测评指令遵循

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。