arXiv:2506.00319cs.CL2025-06ACL被引 2

用树状结构分析大模型能力短板,提升评估与改进效率

SkillVerse : Assessing and Enhancing LLMs with Tree Evaluation

  • 以大模型为裁判,自动生成层级化能力诊断树
  • 少样本学习效果提升25%,弱项预测准确率提高22%
  • 适合需要精细评估和优化模型能力的研究者

随着语言模型应对复杂多面任务的能力提升,其评估方法也需随之演进。精确掌握模型在特定技能上的表现,有助于研究人员制定更科学的开发策略。本文提出SkillVerse,一种无需标注数据的树状结构诊断框架,利用大模型作为评判者,对模型输出进行批判并组织成层次化结构(称为树状图)。该框架可灵活提供任意粒度的能力洞察。我们验证了其在两项下游任务中的有效性:1)通过树搜索算法选择更具信息量的少量示例,使模型上下文学习能力提升25%;2)能以55%的成功率准确预测新模型的弱点,比无SkillVerse时高出22%。

原文摘要 · Abstract (English)

As language models evolve to tackle complex, multifaceted tasks, their evaluation must adapt to capture this intricacy. A granular, skill-specific understanding of model capabilities can empower researchers to make informed model development plans. In this paper, we introduce SkillVerse, an unsupervised tree-structured diagnosis framework for understanding model proficiency in specific abilities. With LLM as a judge, SkillVerse first critiques the model responses, and then organizes them into a hierarchical structure termed dendrogram. Given proficiency at arbitrary levels of granularity, SkillVerse is flexible to produce insights of behaviors of modern large models. We also demonstrate its efficacy in two downstream tasks: 1) improving model in-context learning by 25% using a tree-search algorithm to select more informative few-shot demonstrations, and 2) accurately predicting new model weaknesses with a 55% success rate, 22% higher than without SkillVerse.

模型评估能力诊断少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。