构建统一评估框架与高质量任务集,推动大模型智能体评测标准化。
Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation

- 开发80+基准适配器,实现跨平台智能体统一评估。
- 在54个基准上测试8个模型,最高通过率仅28.0%。
- 推出82个高难度、多样化任务集,适合严肃评测使用。
评估智能体在日益增多的基准任务中表现面临挑战,因多数需复杂环境与集成。本文提出统一评估基础设施Harbor Adapters,支持超过80个基准的通用评估,并通过代码审查与一致性实验验证其可靠性。进一步,在54个基准上对8个不同能力层级的模型进行大规模评测,每个模型均使用Terminus-2及3种原生工具包运行,实现更全面的能力与失败模式分析。同时,构建了Harbor-Index——一个经难度筛选、人工智能与人工审核、审计修复循环优化的82项高质任务集合,覆盖29个基准,保留评估广度与挑战性,且所有配置通过率均不超过30%,最强模型(GPT-5.5 with Codex)最高达28.0%。相关适配器、结果、分析及数据集已开源,助力更可靠、全面的模型智能体评估。
原文摘要 · Abstract (English)
Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a unified evaluation infrastructure for agentic benchmarks. Our work makes three contributions. First, we develop benchmark adapters that port more than 80 benchmarks to evaluate arbitrary agents, and validate them through rigorous code review and parity experiments. Second, we conduct a large-scale evaluation of 8 models spanning capability tiers across 54 benchmarks; every model is run with Terminus-2 and with one of 3 native harnesses. This enables a broader analysis of agent capabilities and failure modes than was previously possible. Third, we introduce Harbor-Index, a curated set of 82 difficult, diverse, and high-quality tasks spanning 29 benchmarks, refined from the adapted suite through difficulty filtering, AI and human audit, and an audit-and-fix loop. Harbor-Index preserves the challenge and breadth of large-scale agentic evaluations while being affordable to run; no evaluated model-harness configuration exceeds 30% pass rate, and the strongest (GPT-5.5 with Codex) reaches 28.0%. We release the adapters, evaluation results, in-depth analysis, and Harbor-Index as open-source artifacts to support more reliable and comprehensive evaluation of language-model agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。