arXiv:2409.19951cs.AIcs.CL2024-09被引 13

发现大模型在跨能力任务中受最弱环节制约,表现远低于预期。

Law of the Weakest Link: Cross Capabilities of Large Language Models

  • 构建7类核心能力与7类跨能力组合,形成系统评估框架。
  • 17个模型在58项跨能力测试中,38项低于所有单项能力。
  • 强调提升最弱能力是优化复杂任务性能的关键突破口。

大型语言模型(LLMs)的发展与评估长期聚焦于单一能力,却忽视了真实任务中多领域知识交叉的需求,我们称之为跨能力。为此,我们定义了七种核心个体能力,并配对生成七种常见跨能力,每种均配有手工构建的分类体系。基于此,提出CrossEval基准,包含1,400个由人工标注的提示,每类能力100个。为确保评估可靠性,邀请专家对4,200条模型输出进行评分,获得8,400条带详细解释的人类评价作为参考。结果表明,在静态评估及能力增强尝试中,当前模型普遍呈现“最弱环节定律”:跨能力表现严重受限于最弱的子能力。在17个模型的58项跨能力得分中,38项低于所有单项能力,20项介于强弱之间但更接近较弱能力。这揭示了模型在跨能力任务中的显著短板,识别并改进最弱能力成为未来研究优化多维复杂场景性能的重中之重。

原文摘要 · Abstract (English)

The development and evaluation of Large Language Models (LLMs) have largely focused on individual capabilities. However, this overlooks the intersection of multiple abilities across different types of expertise that are often required for real-world tasks, which we term cross capabilities. To systematically explore this concept, we first define seven core individual capabilities and then pair them to form seven common cross capabilities, each supported by a manually constructed taxonomy. Building on these definitions, we introduce CrossEval, a benchmark comprising 1,400 human-annotated prompts, with 100 prompts for each individual and cross capability. To ensure reliable evaluation, we involve expert annotators to assess 4,200 model responses, gathering 8,400 human ratings with detailed explanations to serve as reference examples. Our findings reveal that, in both static evaluations and attempts to enhance specific abilities, current LLMs consistently exhibit the "Law of the Weakest Link," where cross-capability performance is significantly constrained by the weakest component. Specifically, across 58 cross-capability scores from 17 models, 38 scores are lower than all individual capabilities, while 20 fall between strong and weak, but closer to the weaker ability. These results highlight the under-performance of LLMs in cross-capability tasks, making the identification and improvement of the weakest capabilities a critical priority for future research to optimize performance in complex, multi-dimensional scenarios.

大模型跨能力评估基准最弱环节

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。