arXiv:2506.08136cs.CL2025-06被引 7

构建真实网页上的经济任务基准,测试智能体的多模态决策能力

EconWebArena: Benchmarking Autonomous Agents on Economic Tasks in Realistic Web Environments

  • 基于82个权威网站设计360个真实经济任务,覆盖多领域复杂流程
  • 多模态大模型在网页导航与视觉信息理解上表现不足,平均准确率仅41%
  • 适合研究经济智能体、多模态推理和网页交互的学者与开发者

我们提出EconWebArena,一个用于评估自主智能体在真实网络环境中完成复杂多模态经济任务的基准。该基准包含来自82个权威网站的360个精心设计的任务,涵盖宏观经济学、劳动力、金融、贸易和公共政策等领域。每个任务要求智能体在实时网站上导航,解读结构化与视觉内容,与真实界面交互,并通过多步流程提取精确且时效性强的数据。基准通过多个大语言模型生成候选任务,再经严格人工筛选确保任务清晰性、可行性与数据源可靠性。相比以往工作,EconWebArena更强调对权威数据源的忠实度与基于网页的经济推理能力。我们评估了多种前沿多模态大模型作为网页智能体的表现,分析失败案例并进行消融实验,检验视觉定位、计划推理与交互设计的影响。结果揭示出显著性能差距,凸显在认知锚定、导航与多模态理解方面的持续挑战,确立EconWebArena作为经济网络智能的严谨评测平台。

原文摘要 · Abstract (English)

We introduce EconWebArena, a benchmark for evaluating autonomous agents on complex, multimodal economic tasks in realistic web environments. The benchmark comprises 360 curated tasks from 82 authoritative websites spanning domains such as macroeconomics, labor, finance, trade, and public policy. Each task challenges agents to navigate live websites, interpret structured and visual content, interact with real interfaces, and extract precise, time-sensitive data through multi-step workflows. We construct the benchmark by prompting multiple large language models (LLMs) to generate candidate tasks, followed by rigorous human curation to ensure clarity, feasibility, and source reliability. Unlike prior work, EconWebArena emphasizes fidelity to authoritative data sources and the need for grounded web-based economic reasoning. We evaluate a diverse set of state-of-the-art multimodal LLMs as web agents, analyze failure cases, and conduct ablation studies to assess the impact of visual grounding, plan-based reasoning, and interaction design. Our results reveal substantial performance gaps and highlight persistent challenges in grounding, navigation, and multimodal understanding, positioning EconWebArena as a rigorous testbed for economic web intelligence.

智能体经济任务网页交互多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。