通过分步引导任务设计,精准定位大模型的组合推理短板。
STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs

- 基于支架式设计,逐步生成控制变量的任务变体。
- 在三个推理基准上发现六款模型的多处特定能力缺失。
- 适合研究模型缺陷、改进评测方法的开发者和研究人员。
评估基准常被用来衡量大语言模型在不同领域的表现,但整体得分难以揭示模型在组合推理能力上的具体短板及其改进方向。为此,我们提出分步引导任务设计(STaD)框架,通过结构化、渐进式支持生成基准任务的可控变体。该方法不逐个分析失败案例,而是系统性地识别模型所缺乏的具体推理技能组合。在六款不同规模的模型上进行实验,结果显示在三个推理基准中存在多个失败点,并揭示了各模型独特的技能差距。
原文摘要 · Abstract (English)
Benchmarks are often used as a standard to understand LLM capabilities in different domains. However, aggregate benchmark scores provide limited insight into compositional skill gaps of LLMs and how to improve them. To make these weaknesses visible, we propose Scaffolded Task Design (STaD) framework. STaD generates controlled variations of benchmark tasks based on the concept of scaffolding, which introduces structured, incremental support in a step-by-step manner. Rather than inspecting failures individually, this approach enables systematic and scalable probing of model behavior by identifying the specific reasoning skill compositions they lack. Treating the LLM as a black box, our experiments on six models of varying sizes reveal multiple failure points in three reasoning benchmarks and highlight each model's unique and distinct skill gaps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。