arXiv:2503.08893cs.CLcs.LG2025-03被引 30

用能力树精准定位大模型短板,指导数据优化。

EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees

  • 构建自然语言描述的能力树,关联测试实例
  • 在MATH和WildChat上比基线更准更全地发现弱点
  • 可指导数据收集,适合模型优化与评测研究者

理想的模型评估应能识别模型缺陷并提供改进方向。针对语言模型(LM)评估,我们提出生成弱点剖析(weakness profile)的问题:基于模型在基准测试中每个实例的表现,生成一组以自然语言表达的弱点。我们设计了一套量化评估方法,用于比较不同弱点剖析方法。引入新方法EvalTree:构建能力树,每个节点代表一个自然语言描述的能力,并链接到评估该能力的测试实例子集;通过提取模型表现差的节点,生成弱点剖析。在MATH和WildChat基准上,EvalTree比基线方法更精确、全面地识别出弱点。弱点剖析还能指导数据收集,由EvalTree引导的数据收集策略比其他方法更有效提升模型性能。我们还揭示了Chatbot Arena的人类投票评估存在的缺陷。为促进后续研究,我们提供了交互式界面,供从业者探索EvalTree构建的能力树。

原文摘要 · Abstract (English)

An ideal model evaluation should achieve two goals: identifying where the model fails and providing actionable improvement guidance. Toward these goals for language model (LM) evaluations, we formulate the problem of generating a weakness profile, a set of weaknesses expressed in natural language, given an LM's performance on every individual instance in a benchmark. We introduce a suite of quantitative assessments to compare different weakness profiling methods. We also introduce a weakness profiling method EvalTree. EvalTree constructs a capability tree where each node represents a capability described in natural language and is linked to a subset of benchmark instances that specifically evaluate this capability; it then extracts nodes where the LM performs poorly to generate a weakness profile. On the MATH and WildChat benchmarks, we show that EvalTree outperforms baseline weakness profiling methods by identifying weaknesses more precisely and comprehensively. Weakness profiling further enables weakness-guided data collection, and training data collection guided by EvalTree-identified weaknesses improves LM performance more than other data collection strategies. We also show how EvalTree exposes flaws in Chatbot Arena's human-voter-based evaluation practice. To facilitate future work, we provide an interface that allows practitioners to interactively explore the capability trees built by EvalTree.

模型评估能力树数据优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。