arXiv:2409.14913cs.CLcs.IR2024-09被引 3

首个评估大模型在真实网页上完成高价值研究任务的基准测试

Towards a Realistic Long-Term Benchmark for Open-Web Research Agents

  • 构建真实金融咨询场景的开放网络研究任务评测体系
  • Claude-3.5 Sonnet和o1-preview表现最佳,显著优于GPT-4o
  • 分步委托子任务的ReAct架构效果最优,适合企业级智能助手研发

我们呈现了针对具有经济价值的白领任务中大语言模型代理性能评估的初步结果。评估聚焦于金融与咨询领域常见的、真实的‘混乱’开放式网络研究任务。通过此工作,我们为大语言模型代理建立了一个评价体系,优秀表现可直接转化为重大经济与社会影响。我们使用o1-preview、GPT-4o、Claude-3.5 Sonnet、Llama 3.1 (405b) 和 GPT-4o-mini 测试了多种代理架构。平均而言,基于Claude-3.5 Sonnet和o1-preview的代理显著优于基于GPT-4o的代理,而基于Llama 3.1 (405b)和GPT-4o-mini的代理则明显落后。在所有大模型中,具备将子任务委派给子代理能力的ReAct架构表现最佳。除量化评估外,还通过分析代理行为轨迹与观察记录进行了定性评估。本研究是首次对代理在真实开放网络上执行复杂、高经济价值分析师式研究能力的深入评估。

原文摘要 · Abstract (English)

We present initial results of a forthcoming benchmark for evaluating LLM agents on white-collar tasks of economic value. We evaluate agents on real-world "messy" open-web research tasks of the type that are routine in finance and consulting. In doing so, we lay the groundwork for an LLM agent evaluation suite where good performance directly corresponds to a large economic and societal impact. We built and tested several agent architectures with o1-preview, GPT-4o, Claude-3.5 Sonnet, Llama 3.1 (405b), and GPT-4o-mini. On average, LLM agents powered by Claude-3.5 Sonnet and o1-preview substantially outperformed agents using GPT-4o, with agents based on Llama 3.1 (405b) and GPT-4o-mini lagging noticeably behind. Across LLMs, a ReAct architecture with the ability to delegate subtasks to subagents performed best. In addition to quantitative evaluations, we qualitatively assessed the performance of the LLM agents by inspecting their traces and reflecting on their observations. Our evaluation represents the first in-depth assessment of agents' abilities to conduct challenging, economically valuable analyst-style research on the real open web.

大模型代理开放网络研究评测经济价值

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。