arXiv:2410.13826cs.LGcs.AI2024-10ICLR被引 10

通过分析模型推理过程,揭示大模型在不同技能上的优劣对比。

Unearthing Skill-Level Insights for Understanding Trade-Offs of Foundation Models

  • 从模型推理文本中自动提取隐含技能
  • 发现Gemini比GPT-4o在计算摩尔质量上高18%,但在宪法法律应用上低19%
  • 可指导模型路由,提升整体准确率3%

随着模型能力增强,评估变得越来越复杂,单一基准测试甚至同一实例会同时考察多种技能。然而,仅看整体准确率会掩盖各技能表现,未能充分利用现代基准中的丰富信号。本文提出一种自动方法,通过分析模型生成的推理过程来识别评估实例背后的潜在技能。在12个基准上对4.6万实例进行技能解析与推断后,发现许多技能在不同基准中普遍存在,由此构建了数百个技能切片(即测试同一技能的实例集合)。对这些切片的准确率分析揭示了新的模型权衡关系:例如,相比GPT-4o和Claude 3.5 Sonnet,Gemini 1.5 Pro在“计算摩尔质量”上平均高出18%,但在“应用宪法法”上低19%,尽管三者整体准确率差异仅为0.4%。此外,我们证明该方法具有实际价值——将每个实例路由到相关技能最强的模型,可在12个数据集上实现3%的准确率提升。本研究提出的技能切片与框架为模型评估开辟新路径,通过技能级分析实现更细粒度、可操作的能力理解。

原文摘要 · Abstract (English)

With models getting stronger, evaluations have grown more complex, testing multiple skills in one benchmark and even in the same instance at once. However, skill-wise performance is obscured when inspecting aggregate accuracy, under-utilizing the rich signal modern benchmarks contain. We propose an automatic approach to recover the underlying skills relevant for any evaluation instance, by way of inspecting model-generated rationales. After validating the relevance of rationale-parsed skills and inferring skills for $46$k instances over $12$ benchmarks, we observe many skills to be common across benchmarks, resulting in the curation of hundreds of skill-slices (i.e. sets of instances testing a common skill). Inspecting accuracy over these slices yields novel insights on model trade-offs: e.g., compared to GPT-4o and Claude 3.5 Sonnet, on average, Gemini 1.5 Pro is $18\%$ more accurate in "computing molar mass", but $19\%$ less accurate in "applying constitutional law", despite the overall accuracies of the three models differing by a mere $0.4\%$. Furthermore, we demonstrate the practical utility of our approach by showing that insights derived from skill slice analysis can generalize to held-out instances: when routing each instance to the model strongest on the relevant skills, we see a $3\%$ accuracy improvement over our $12$ dataset corpus. Our skill-slices and framework open a new avenue in model evaluation, leveraging skill-specific analyses to unlock a more granular and actionable understanding of model capabilities.

模型评估技能分析大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。