首个评估大模型理解办公文件能力的公开基准
Office Comprehension Benchmark

- 构建双轨评测体系,涵盖文件结构感知与跨文档专业推理
- 最强模型在专业领域问答中仅达59.3%准确率
- 适合评估办公场景下大模型的真实应用能力
我们提出Office Comprehension Bench(OCB),首个针对大模型在原生文件格式(.docx, .xlsx, .pptx)及其变体上对Word、Excel和PowerPoint理解能力的公开评测基准。OCB包含两个赛道:文件保真度问答测试对表格、图表、嵌入图像、公式及页眉、演讲者备注、命名区域等元素的结构与视觉感知;领域问答测试则在12个专业领域内评估基于真实行业文档的专家级推理能力,要求多步分析与跨文档综合。每个参考答案被分解为原子级可二元评分的陈述,由多个大模型裁判独立打分。即使最强前沿系统在默认推理模式下,领域问答得分也仅约59.3%;增加思维深度未显著提升性能,而切换至更高产品层级仅带来小幅改善。我们公开数据集、评估工具、裁判提示及公共排行榜。
原文摘要 · Abstract (English)
We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint comprehension over native file formats (.docx, .xlsx, .pptx) and their variants. OCB consists of two tracks. File Fidelity Q&A tests structural and visual perception of office artifacts - tables, charts, embedded images, formulas, and app-specific elements such as headers, speaker notes, and named ranges. Domain Q&A tests expert-level reasoning grounded in real-world industry documents across 12 professional domains, with queries requiring multi-step analysis and synthesis across documents. Each reference answer is decomposed into atomic, binary-gradable claims, and an ensemble of LLM judges scores responses against each claim independently. Even the strongest frontier system in its default reasoning mode reaches only about 59.3% on Domain Q&A; increasing thinking depth within a tier does not move performance materially, while moving to a higher product tier yields modest gains. We release the dataset, evaluation tooling, judge prompt, and a public leaderboard.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。