不同思维框架让同一模型表现差28个百分点,说明评测结果受工具链影响极大。
Scaffold Effects on GAIA: A Controlled Comparison

- 固定任务条件,对比三种思维框架在五模型上的表现差异。
- 同一模型用不同框架,准确率最高差28个百分点,最小差距超10点。
- 越强模型越依赖结构化框架,且跨厂商模型无明显多智能体优势。
现有代理能力评分混淆了模型自身能力与支架设计的影响,且这种诱发差距在受控条件下未被充分量化。本研究在预注册下,对三个支架(ReAct、规划-执行-评价多智能体、规划后执行)在五个模型(Claude Opus 4.7、Sonnet 4.6、Haiku 4.5;Gemini 3.1 Pro Preview;GPT-5.5)上,于GAIA验证集第1、2级进行受控对比,每题尝试三次。支架选择本身可使单个模型(Opus,Level 2,稳健子集)的准确率波动达28个百分点,证实预注册假设:支架差异导致至少10分差距。预注册预测‘更强大模型对支架不敏感’被推翻:支架效应在每个数据子集中均显著,但最强大的Anthropic模型在更高难度下从结构化支架中获益最多,且层级提升仅在Level 1的稳健子集中成立。多智能体优势仅在Anthropic家族内显现,跨厂商模型无此现象,表明模型族而非能力层级是关键变量;预估的规划-执行优势在文件读取任务上被证伪。结构化支架减少工具调用次数,却能更好恢复中途错误,在更难任务中表现更优;单一组合(Gemini + 规划后执行)在两级均成本最低且在Level 2最准确。结果表明,单支架能力评分是支架依赖的估计值,模型进步并不保证诱发差距缩小。
原文摘要 · Abstract (English)
Published agent capability scores conflate what a model can do with what its scaffold lets it do, and the magnitude of this elicitation gap is not well characterized under controlled conditions. This study executes a pre-registered controlled comparison of three scaffolds (ReAct, a Planner-Actor-Rater multi-agent design, and planner-then-executor) across five models from three providers (Claude Opus 4.7, Sonnet 4.6, Haiku 4.5; Gemini 3.1 Pro Preview; GPT-5.5) on GAIA validation Levels 1 and 2, holding tasks and conditions fixed, with three attempts per question. Scaffold choice alone moves measured accuracy by as much as 28 percentage points within a single model (Opus, Level 2, robust slice), confirming the pre-registered hypothesis that scaffold variation produces gaps of at least 10 points. The pre-registered prediction that more capable models would be less scaffold-sensitive is rejected in direction: scaffold effects vary significantly by model in every dataset slice, but the most capable Anthropic model gains the most from structured scaffolds at the harder level, and tier-scaling holds only at Level 1 under the robust slice. The multi-agent advantage over ReAct at Level 2 appears within the Anthropic family but not for the cross-provider models, making model family rather than capability tier the conditioning variable, and the predicted planner-executor advantage on file-reading tasks is falsified. Structured scaffolds make fewer tool calls yet recover more often from mid-trajectory errors at the harder level, and a single cell (Gemini with planner-then-executor) is the cheapest at both levels and the most accurate at Level 2. These results indicate that single-scaffold capability numbers are scaffold-conditional estimates and that the elicitation gap is not guaranteed to shrink as models improve.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。