arXiv:2606.10956cs.AIcs.CL2026-06

用国家办公考试测试大模型,发现它们自动化办公能力仍不足。

Mind the Gap: Can Frontier LLMs Pass a Standardized Office Proficiency Exam?

论文配图:Mind the Gap: Can Frontier LLMs Pass a Standardized Office Proficiency Exam?
图 1 · 摘自论文原文
  • 用中国计算机等级考试题评估大模型办公自动化能力
  • 最强模型仅达68.8%得分,远低于人类参考分95.5%
  • 暴露现有大模型在精细操作与多软件协同上的短板

大型语言模型(LLM)代理在计算机自动化中的应用日益加速,但其在复杂专业级生产力软件中的表现尚未充分验证。本文认为办公自动化是评估文档自动化能力的理想场景,因其需要长程规划、精确参数配置及多应用集成。为此,我们基于中国计算机等级考试(NCRE),设计包含Word、Excel和PowerPoint的200个综合实操任务,采用7,118项可机器评分的标准进行100分制评分,以得分率(Score Rate, SR)衡量整体表现。我们评估了7个前沿LLM,发现单轮对话模型最高仅得36.6%。即便采用带执行反馈、迭代修复和更广办公访问权限的智能体系统,最高得分也仅为68.8%,仍显著低于95.5%的人类参考得分。实验表明,尽管代码生成能力提升,当前代码生成型大模型与智能体系统在实现可靠细粒度办公文档自动化方面仍面临重大挑战。

原文摘要 · Abstract (English)

The deployment of Large Language Model (LLM) agents for computer automation is accelerating, yet their ability to navigate complex, professional-grade productivity software is largely untested. We argue that Office automation is an ideal environment for benchmarking document-automation capability, as it requires long-horizon planning and reasoning, precise parameter configuration, and multi-application integration. To quantify this capability, we introduce an evaluation based on China's National Computer Rank Examination (NCRE), featuring 200 comprehensive practical-operation tasks across Word, Excel, and PowerPoint. Each task is scored on a 100-point rubric scale using 7,118 machine-gradable criteria, and Score Rate (SR) denotes the mean percentage of rubric points earned across these tasks. We benchmark 7 frontier LLMs and observe stark limitations: single-turn models score a maximum of 36.6%. A stronger agentic system with execution feedback, iterative repair, and broader Office automation access reaches 68.8%, but remains below the 95.5% community-reference score used as a scoring sanity check. Ultimately, our experiments demonstrate that despite recent advancements in code generation, achieving reliable fine-grained Office document automation remains a significant challenge for current code-generating LLM and agent systems.

办公自动化大模型评测智能体NCRE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。