用真人投票评估大模型生成的表格质量,发现好结果因任务而异。
SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks
- 通过盲评对比生成表格,捕捉用户偏好
- 不同任务下优质表格的风格结构差异显著
- 适合研究表格生成与人类偏好的对齐
我们研究端到端表格生成任务,即语言模型根据自然语言描述生成满足显性和隐性约束的电子表格。为此,我们提出SpreadsheetArena平台,通过大模型对生成表格进行盲态成对偏好投票来评估性能。与一般对话或文本生成不同,表格生成具有明确的多维输出结构和复杂的交互与布局需求。我们发现,不同提示下受青睐的表格在风格、结构和功能上存在显著差异。针对金融类提示的专家评估显示,即使排名靠前的模型也未能稳定生成符合领域最佳实践的表格。我们已上线实时评测平台,并发布包含提示、生成表格及偏好投票的数据集,旨在推动对表格这类复杂开放任务中大模型表现的研究。
原文摘要 · Abstract (English)
We consider the task of end-to-end spreadsheet generation, where language models produce spreadsheet artifacts to satisfy users' explicit and implicit constraints, specified in natural language. We introduce SpreadsheetArena, a platform for evaluating models' performance on the task via blind pairwise preference votes of LLM-generated spreadsheet workbooks. As with other complex, open-ended tasks, relevant evaluation criteria can vary greatly across use cases, often in ways that are difficult to formalize. Compared to general dialogue or text generation settings, spreadsheet generation presents unique challenges and opportunities: the task output structure is well-defined and multi-dimensional, and there are often complex interactivity and layout considerations. We observe that stylistic, structural, and functional features of preferred spreadsheets vary meaningfully across prompts. Expert evaluations of spreadsheets for finance prompts suggest that even highly ranked models do not reliably produce spreadsheets aligned with domain-specific best practices. We host a live arena and release a dataset of prompts, generated spreadsheets, and preference votes, which we hope will facilitate further study of tasks operating over spreadsheets as a challenging and interesting class of complex, open-ended tasks for LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。