arXiv:2608.23475cs.AI2026-08

测试大模型能否像人一样从例子中提炼任务规则并应用。

StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models

论文配图:StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models
图 1 · 摘自论文原文
  • 从BIG-Bench中筛选可提炼策略的任务,构建参考规则
  • 发现不同任务类别策略效用差异显著,生成与执行均影响效果
  • 适合研究模型少样本学习能力或提示工程的学者

随着大语言模型在数据稀缺和动态任务场景中的广泛应用,少样本上下文学习(ICL)已成为关键任务适应范式。然而,直接ICL通常仅使用少量示例而未显式抽象任务规则,对示例构造敏感。相比之下,人类学习者常先从示例中总结任务规则,再应用于新实例。为评估此能力,我们提出StrategyBench,从BIG-Bench中选取可诱导策略的任务,构建参考策略,并定义沿两个维度的评估指标:策略质量和下游实用性。我们进一步从任务变化、模型配置和适应设置三方面分析策略归纳,涵盖类别差异、生成-执行组合、示范设计及基于SFT的适应。实验表明,显式策略的效用在不同任务类别间存在显著差异,且取决于策略生成与执行条件。基准数据集已公开:https://anonymous.4open.science/r/StrategyBench-D53C。

原文摘要 · Abstract (English)

As large language models are increasingly used in data-scarce and evolving task scenarios, few-shot in-context learning (ICL) has become a key paradigm for task adaptation. However, direct ICL often uses a small set of examples without explicitly abstracting task rules, making it sensitive to example construction. In contrast, human learners often reduce such sensitivity by first summarizing task rules from examples and then applying them to new instances. To evaluate this ability, we propose StrategyBench, which selects strategy-inducible tasks from BIG-Bench, constructs reference strategies, and defines evaluation metrics along two dimensions: strategy quality and downstream utility. We further analyze strategy induction from three perspectives: task variation, model configuration, and adaptation setting, covering category-wise differences, generator-executor choices, demonstration design, and SFT-based adaptation. Experiments show that explicit strategy utility differs substantially across task categories and depends on both strategy generation and execution conditions. The benchmark is released at: https://anonymous.4open.science/r/StrategyBench-D53C.

大模型少样本学习策略归纳评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。