构建工业表格生成报告的基准,评估大模型真实场景文本生成能力
T2R-bench: A Benchmark for Generating Article-Level Reports from Real World Industrial Tables
- 提出表格转报告任务,设计多领域真实工业表格数据集
- 25个主流大模型平均得分仅62.71,顶尖模型仍有提升空间
- 适合关注工业AI、文档生成与评测标准的研究者
大量研究探索了大语言模型(LLMs)在表格推理中的能力,但将表格信息转化为报告这一关键任务在工业应用中仍面临巨大挑战。该任务存在两大问题:一是表格结构复杂多样导致推理效果不佳;二是现有表格基准无法有效评估实际应用表现。为此,我们提出表格转报告任务,并构建名为T2R-bench的双语基准,涵盖457个来自真实工业场景的表格,覆盖19个行业领域及4种表格类型。我们还设计了一套公平的报告生成质量评估标准。在25个广泛使用的LLMs上进行实验表明,即使是最先进的模型Deepseek-R1,整体得分也仅为62.71,说明当前大模型在此任务上仍有显著改进空间。
原文摘要 · Abstract (English)
Extensive research has been conducted to explore the capabilities of large language models (LLMs) in table reasoning. However, the essential task of transforming tables information into reports remains a significant challenge for industrial applications. This task is plagued by two critical issues: 1) the complexity and diversity of tables lead to suboptimal reasoning outcomes; and 2) existing table benchmarks lack the capacity to adequately assess the practical application of this task. To fill this gap, we propose the table-to-report task and construct a bilingual benchmark named T2R-bench, where the key information flow from the tables to the reports for this task. The benchmark comprises 457 industrial tables, all derived from real-world scenarios and encompassing 19 industry domains as well as 4 types of industrial tables. Furthermore, we propose an evaluation criteria to fairly measure the quality of report generation. The experiments on 25 widely-used LLMs reveal that even state-of-the-art models like Deepseek-R1 only achieves performance with 62.71 overall score, indicating that LLMs still have room for improvement on T2R-bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。