构建可验证的文档操作评估框架,揭示智能体在复杂文档任务中的关键缺陷。
DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations

- 按真实工作流拆解文档操作为原子层级,构建分层评估体系。
- 顶尖智能体在长程耦合任务中仍严重依赖状态追踪和语义理解。
- 发现三类核心失败模式:状态丢失、浅层校验、元数据误删,适合研究智能体可靠性者关注。
随着自主智能体迅速发展,其可靠操作通用数字文档的能力已成为实现通用AI助手和自动化复杂工作流的关键。本文提出DocOps,一个基于分层分类体系的确定性可验证评估框架,将源于真实场景的文档操作分解为原子维度与递增复杂度的工作流。基于此框架,我们系统评估了多种闭源与开源模型在不同智能体架构下的表现,发现即使最先进的前沿配置在处理高度耦合的长程任务时仍存在显著局限。进一步的细粒度分析揭示现有智能体的三种主要失效模式:长期状态追踪崩溃、浅层语义验证不足以及对结构元数据的破坏性编辑。最终,本工作揭示了智能体在维持全局文档一致性方面的能力边界,为未来设计稳健、非破坏性的复杂数字生态系统代理提供了重要启示。
原文摘要 · Abstract (English)
As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows. In this paper, we introduce DocOps, a deterministically verifiable evaluation framework underpinned by a hierarchical taxonomy that deconstructs document operations inspired by real-world practices into atomic dimensions and escalating workflow complexities. Based on DocOps, we systematically evaluate representative closed- and open-source models across various agentic harnesses, revealing that even the most advanced frontier configurations still exhibit profound limitations when handling highly coupled, long-range tasks. Furthermore, a fine-grained analysis of existing agents' manipulation behaviors uncovers 3 key failure modes: long-term state tracking collapse, shallow semantic verification, and destructive editing of structural metadata. Ultimately, our work exposes the capability boundaries of agents in maintaining global document consistency, shedding light on the future design of robust, non-destructive agents for complex digital ecosystems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。