评测顶尖研究型智能体在咨询任务中的表现,发现均未达标且各有致命缺陷。
Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps
- 设计70个专家撰写的咨询题,含认知陷阱,测试多文档决策级输出能力。
- 三款模型平均通过率仅13%-16%,唯一领先的是o3,但仍有22.9分差距。
- 首次融合验证器与评分卡,适合评估企业级智能体的真实可用性。
前沿深度研究智能体正快速部署于企业流程,却缺乏有效评估。现有基准仅测事实回忆、单跳问答或通用代理技能,无法反映智能体需完成的多文档、决策级交付成果。本文提出一个包含70个领域专家撰写的管理咨询题的基准,每题嵌入认知陷阱以惩罚表面模式推理。评估采用双重维度:确定性二元验证器(每任务平均14.9个)和五项制评分卡(数据完整性、分析严谨性、相关性与聚焦度、执行精度、格式与可交付性),合并为验证-评分分(VRS,0–100)。在联合阈值下(评分均值≥2.5且验证通过率≥80%),三款模型接受率均极低:o3为15.7%,Claude为12.9%,Gemini为12.9%。配对差异无统计显著性。连续VRS中,o3领先(61.4 [95% CI: 55.2, 67.5]),其次Gemini(52.6),Claude最低(38.5);o3与Claude差距达22.9分(p<0.001,Bonferroni校正后仍显著)。无一模型平均评分超2.0的“合格”线,亦无一模型验证通过率达80%门槛。各模型失败方式不同:Claude在数据捏造和文件访问上失误最多;o3传播连锁计算错误;Gemini在完美通过率与灾难性崩溃间波动。该基准、评估代码与完整提示语料已公开发布。
原文摘要 · Abstract (English)
Frontier deep research agents (DRAs) are being deployed in enterprise workflows faster than they are being evaluated. Existing benchmarks measure factual recall, single-hop QA, or generic agentic skill, and miss the multi-document, decision-grade deliverables DRAs are asked to produce. We introduce a benchmark of 70 SME-authored management consulting prompts, each embedding cognitive traps that penalize surface-pattern reasoning. Three frontier agents, namely Claude Opus~4.6, OpenAI o3-deep-research and Gemini~3.1~Pro deep-research, are scored on two complementary layers: deterministic binary verifiers (mean 14.9 per task) and a five-criterion 0--3 SME rubric (Data Integrity, Analytical Rigor, Relevance \& Focus, Execution Precision, Format \& Deliverability), combined into a Verifier-Rubric Score (VRS, 0--100). Acceptance under a joint threshold (rubric mean $\geq 2.5$ and verifier pass rate $\geq 80\%$) is uniformly low: o3 15.7\%, Claude 12.9\%, Gemini 12.9\%. Pairwise differences are statistically indistinguishable. On the continuous VRS, o3 leads (61.4~[CI: 55.2,\,67.5]), followed by Gemini (52.6) and Claude (38.5); the o3--Claude gap ($Δ{=}22.9$, $p{<}0.001$) survives Bonferroni correction. No agent averages above the rubric's ``adequate'' threshold of 2.0; no agent's mean verifier pass rate reaches the 80\% acceptance floor. Each agent fails distinctively: Claude leads on data fabrication and file-access failures; o3 propagates cascading computation errors; Gemini oscillates between the highest perfect-verifier rate and the most catastrophic collapses. The benchmark, evaluation code, and full prompt corpus are publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。