arXiv:2510.19864cs.SEcs.CL2025-10被引 2

用大模型自动把表格操作转成人类可读说明,提升协作与可维护性。

SODBench: A Large Language Model Approach to Documenting Spreadsheet Operations

  • 基于111个操作代码生成自然语言描述,构建SOD任务基准
  • GPT-4o等模型在多种指标上表现良好,证明文档生成可行性
  • 适合需要自动化文档、知识传承的财务与办公场景

大量知识工作者在商业、会计和金融领域使用电子表格。然而,缺乏系统性的文档方法阻碍了自动化、协作与知识传递,可能导致重要机构知识流失。本文提出表格操作文档(SOD)任务,即从表格操作中生成可读的人类语言解释。尽管先前研究已利用大语言模型生成表格操作代码,但将代码转化为自然语言用于SOD仍较少被探索。为此,我们构建了一个包含111个表格操作代码片段及其对应自然语言摘要的基准数据集。评估了五种LLM:GPT-4o、GPT-4o-mini、LLaMA-3.3-70B、Mixtral-8x7B和Gemma2-9B,采用BLEU、GLEU、ROUGE-L和METEOR指标。结果表明,大模型能生成准确的表格文档,使SOD成为提升可复现性、可维护性和协作流程的可行前置步骤,尽管仍存在需解决的挑战。

原文摘要 · Abstract (English)

Numerous knowledge workers utilize spreadsheets in business, accounting, and finance. However, a lack of systematic documentation methods for spreadsheets hinders automation, collaboration, and knowledge transfer, which risks the loss of crucial institutional knowledge. This paper introduces Spreadsheet Operations Documentation (SOD), an AI task that involves generating human-readable explanations from spreadsheet operations. Many previous studies have utilized Large Language Models (LLMs) for generating spreadsheet manipulation code; however, translating that code into natural language for SOD is a less-explored area. To address this, we present a benchmark of 111 spreadsheet manipulation code snippets, each paired with a corresponding natural language summary. We evaluate five LLMs, GPT-4o, GPT-4o-mini, LLaMA-3.3-70B, Mixtral-8x7B, and Gemma2-9B, using BLEU, GLEU, ROUGE-L, and METEOR metrics. Our findings suggest that LLMs can generate accurate spreadsheet documentation, making SOD a feasible prerequisite step toward enhancing reproducibility, maintainability, and collaborative workflows in spreadsheets, although there are challenges that need to be addressed.

大模型表格文档自动化知识传承

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。