评测大模型代理在复杂办公任务中的长期协作能力。
OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows
- 构建多应用长周期任务基准,模拟真实办公场景
- 300+真实任务与302个合成复杂任务并行评估
- 自动生成流程的多智能体框架,支持规模化扩展
由大语言模型驱动的自主代理正广泛应用于需要复杂、长周期工作流的真实场景。然而,现有基准主要关注独立、原子化的任务,无法捕捉现实场景中所需的长期上下文依赖和多轮交互协同。为此,我们提出OdysseyBench,一个涵盖Word、Excel、PDF、邮件和日历等多样办公应用的综合性基准。该基准包含两个互补部分:基于真实用例的OdysseyBench+(300项任务)和新生成的复杂任务集OdysseyBench-Neo(302项任务)。每项任务要求代理从长期交互历史中识别关键信息,并跨多个应用进行多步推理。为实现可扩展的基准生成,我们提出HomerAgents——一个通过系统性环境探索、任务生成与对话合成自动构建长周期工作流基准的多智能体框架。大量实验证明,OdysseyBench能有效挑战当前先进大模型代理,在复杂真实情境下的能力评估上优于传统原子任务基准。我们已开源OdysseyBench与HomerAgents,以推动该方向的研究发展。
原文摘要 · Abstract (English)
Autonomous agents powered by large language models (LLMs) are increasingly deployed in real-world applications requiring complex, long-horizon workflows. However, existing benchmarks predominantly focus on atomic tasks that are self-contained and independent, failing to capture the long-term contextual dependencies and multi-interaction coordination required in realistic scenarios. To address this gap, we introduce OdysseyBench, a comprehensive benchmark for evaluating LLM agents on long-horizon workflows across diverse office applications including Word, Excel, PDF, Email, and Calendar. Our benchmark comprises two complementary splits: OdysseyBench+ with 300 tasks derived from real-world use cases, and OdysseyBench-Neo with 302 newly synthesized complex tasks. Each task requires agent to identify essential information from long-horizon interaction histories and perform multi-step reasoning across various applications. To enable scalable benchmark creation, we propose HomerAgents, a multi-agent framework that automates the generation of long-horizon workflow benchmarks through systematic environment exploration, task generation, and dialogue synthesis. Our extensive evaluation demonstrates that OdysseyBench effectively challenges state-of-the-art LLM agents, providing more accurate assessment of their capabilities in complex, real-world contexts compared to existing atomic task benchmarks. We believe that OdysseyBench will serve as a valuable resource for advancing the development and evaluation of LLM agents in real-world productivity scenarios. In addition, we release OdysseyBench and HomerAgents to foster research along this line.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。