测试GPT生成电子表格模型的可靠性,发现虽有潜力但结果不稳定需人工审核。
Spreadsheet Modeling Experiments Using GPTs on Small Problem Statements and the Wall Task
- 用简单问题测试GPT生成电子表格,评估结构与准确性
- 生成模型常不一致且不可复现,存在信心与流程双重问题
- 适合想快速出草稿的用户,但专业场景仍需人工把关
本文研究基于GPT的工具在构建可复用分析型电子表格模型中的应用。经过筛选,选取pulsrai.com的Excel AI进行详细测试。通过在简单问题陈述上的结构化实验,依据ERFR标准(每个输入在单元格;单元格公式;无硬编码数字;标签;准确)评估其表现。结果显示,尽管Excel AI能生成结构良好模型,但结果不一致且常不可复现。识别出两大核心挑战:'信心问题'与'工作流问题',凸显熟练用户验证和调整GPT生成表格的必要性。虽然GPT在生成草稿模型方面展现潜力,可能缩短开发时间或降低技能门槛,但当前工具仍不足以用于专业场景。文章最后提出未来研究方向:提示工程、可复现性及大规模建模任务。
原文摘要 · Abstract (English)
This paper investigates how GPT-based tools can assist in building reusable analytical spreadsheet models. After a screening, we evaluate five GPT extensions and select Excel AI by pulsrai.com for detailed testing. Through structured experiments on simple problem statements, we assess Excel AI's performance against the ERFR criteria (each input in a cell; cell formulas; no hardwired numbers; labels; accurate). Results show that while Excel AI can produce well-structured models, it is inconsistent and often non-reproducible. We identify two central challenges - "the problem of confidence" and "the problem of workflow" - which highlight the need for skilled users to verify and adapt GPT-generated spreadsheets. Though GPTs show promise for generating draft models that may reduce development time or lower skill requirements, current tools remain unreliable for professional use. We conclude with recommendations for future research into prompt engineering, reproducibility, and larger-scale modeling tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。