arXiv:2606.08529cs.AIcs.CL2026-06

不同思维框架让同一模型表现差28个百分点,说明评测结果受工具链影响极大。

Scaffold Effects on GAIA: A Controlled Comparison

论文配图:Scaffold Effects on GAIA: A Controlled Comparison
图 1 · 摘自论文原文
  • 固定任务条件,对比三种思维框架在五模型上的表现差异。
  • 同一模型用不同框架,准确率最高差28个百分点,最小差距超10点。
  • 越强模型越依赖结构化框架,且跨厂商模型无明显多智能体优势。

现有代理能力评分混淆了模型自身能力与支架设计的影响,且这种诱发差距在受控条件下未被充分量化。本研究在预注册下,对三个支架(ReAct、规划-执行-评价多智能体、规划后执行)在五个模型(Claude Opus 4.7、Sonnet 4.6、Haiku 4.5;Gemini 3.1 Pro Preview;GPT-5.5)上,于GAIA验证集第1、2级进行受控对比,每题尝试三次。支架选择本身可使单个模型(Opus,Level 2,稳健子集)的准确率波动达28个百分点,证实预注册假设:支架差异导致至少10分差距。预注册预测‘更强大模型对支架不敏感’被推翻:支架效应在每个数据子集中均显著,但最强大的Anthropic模型在更高难度下从结构化支架中获益最多,且层级提升仅在Level 1的稳健子集中成立。多智能体优势仅在Anthropic家族内显现,跨厂商模型无此现象,表明模型族而非能力层级是关键变量;预估的规划-执行优势在文件读取任务上被证伪。结构化支架减少工具调用次数,却能更好恢复中途错误,在更难任务中表现更优;单一组合(Gemini + 规划后执行)在两级均成本最低且在Level 2最准确。结果表明,单支架能力评分是支架依赖的估计值,模型进步并不保证诱发差距缩小。

原文摘要 · Abstract (English)

Published agent capability scores conflate what a model can do with what its scaffold lets it do, and the magnitude of this elicitation gap is not well characterized under controlled conditions. This study executes a pre-registered controlled comparison of three scaffolds (ReAct, a Planner-Actor-Rater multi-agent design, and planner-then-executor) across five models from three providers (Claude Opus 4.7, Sonnet 4.6, Haiku 4.5; Gemini 3.1 Pro Preview; GPT-5.5) on GAIA validation Levels 1 and 2, holding tasks and conditions fixed, with three attempts per question. Scaffold choice alone moves measured accuracy by as much as 28 percentage points within a single model (Opus, Level 2, robust slice), confirming the pre-registered hypothesis that scaffold variation produces gaps of at least 10 points. The pre-registered prediction that more capable models would be less scaffold-sensitive is rejected in direction: scaffold effects vary significantly by model in every dataset slice, but the most capable Anthropic model gains the most from structured scaffolds at the harder level, and tier-scaling holds only at Level 1 under the robust slice. The multi-agent advantage over ReAct at Level 2 appears within the Anthropic family but not for the cross-provider models, making model family rather than capability tier the conditioning variable, and the predicted planner-executor advantage on file-reading tasks is falsified. Structured scaffolds make fewer tool calls yet recover more often from mid-trajectory errors at the harder level, and a single cell (Gemini with planner-then-executor) is the cheapest at both levels and the most accurate at Level 2. These results indicate that single-scaffold capability numbers are scaffold-conditional estimates and that the elicitation gap is not guaranteed to shrink as models improve.

代理评估思维框架评测偏差多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。