构建多语言网页交互评测基准,推动全球智能体发展
X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System
- 设计跨语言网页交互评测框架,评估智能体规划与交互能力
- 测试多种大模型在多语言场景下表现,发现顶尖模型仍存短板
- 适合关注多语言智能体、跨语言对齐的研究者与开发者
近年来,基于大语言模型的智能体在交互环境中取得显著进展,引发学术界与产业界广泛关注。然而,当前研究主要集中于英语场景,而全球有超过7000种语言,均需具备相应的智能体服务能力。现有语言智能体在满足多语言应用需求方面仍显不足。为此,我们提出X-WebAgentBench,一个全新的多语言交互式网页评测基准,用于评估语言智能体在多种语言下的规划与交互性能,推动全球智能体智能的发展。我们还评估了多种大模型及跨语言对齐方法的效果,发现即使先进模型如GPT-4o,在结合跨语言技术后仍未能达到理想表现。X-WebAgentBench有望成为真实世界多语言智能体应用的重要评测工具。
原文摘要 · Abstract (English)
Recently, large language model (LLM)-based agents have achieved significant success in interactive environments, attracting significant academic and industrial attention. Despite these advancements, current research predominantly focuses on English scenarios. In reality, there are over 7,000 languages worldwide, all of which demand access to comparable agentic services. Nevertheless, the development of language agents remains inadequate for meeting the diverse requirements of multilingual agentic applications. To fill this gap, we introduce X-WebAgentBench, a novel multilingual agent benchmark in an interactive web environment, which evaluates the planning and interaction performance of language agents across multiple languages, thereby contributing to the advancement of global agent intelligence. Additionally, we assess the performance of various LLMs and cross-lingual alignment methods, examining their effectiveness in enhancing agents. Our findings reveal that even advanced models like GPT-4o, when combined with cross-lingual techniques, fail to achieve satisfactory results. We hope that X-WebAgentBench can serve as a valuable benchmark for multilingual agent scenario in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。