arXiv:2608.07873cs.AI2026-08

用自动化流水线生成表格任务评估集,测试大模型写公式、图表等能力。

Back to the Future: A workbook time machine for spread sheet creation benchmarks

论文配图:Back to the Future: A workbook time machine for spread sheet creation benchmarks
图 1 · 摘自论文原文
  • 构建时间机器流水线,自动生成输入/输出表格与查询三元组。
  • 创建150项任务的wtmbench基准,覆盖4类表格对象和多层指令粒度。
  • 发现查询具体性、代理协作与接口调用方式显著影响模型表现。

我们提出Workbook Time Machine,一个自动化流水线,用于生成评估语言模型在电子表格中创建衍生对象(公式、图表、数据透视表、条件格式)能力的基准。该方法应用于公开工作簿语料库,生成wtmcorpus——一个包含(输入工作簿,输出工作簿,查询)三元组的集合,涵盖四种对象类型且复杂度各异。基于此语料,我们构建了wtmbench,一个包含150个任务的评估基准,其查询覆盖三个层次的具体性。我们在不同对象类型、步骤复杂度和指令粒度下,对现有电子表格操作代理和基线模型进行评估。结果表明,查询具体性、代理编排策略以及控制电子表格的接口API对大模型在Excel任务上的表现有显著影响。

原文摘要 · Abstract (English)

We introduce the workbook time machine, a pipeline that automatically creates benchmarks evaluating the ability of language models to create derived objects in spreadsheets (formulas, charts, pivot tables, and conditional formatting). Applied to public workbook corpora, it produces wtmcorpus--a collection of (input workbook, output workbook, query) triples spanning four artifact types and varying complexity. From this corpus we curate wtmbench, a 150-task evaluation benchmark with queries at three levels of specificity. We evaluate existing spreadsheet manipulation agents and baselines on wtmbench across artifact types, step complexity, and instruction granularity. Our evaluations show that query specificity, agent orchestration, and interface API used to control spreadsheets play a big role in LLM performance on Excel tasks.

表格生成评测基准大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。