arXiv:2507.16271cs.CL2025-07

评测大模型从碎片文档中构建结构化表格的能力

Beyond Isolated Dots: Benchmarking Structured Table Construction as Deep Knowledge Extraction

  • 设计跨领域11项任务,要求模型动态生成适配上下文的表格模式
  • 最先进模型在复杂文档重构中表现仍不理想,准确率低于40%
  • 适合研究知识提取、文档理解与大模型评估的学者使用

随着大语言模型(LLMs)的兴起,人们期望其能从复杂现实文档(如论文、报告)中有效提取显式信息。然而,多数模型生成的是杂乱无章、难以追溯的段落式回答。为弥合这一差距,我们提出了「有组织提取基准」(AOE),一个包含多语言、不同长度数据与文档的新基准,用于系统评估模型理解碎片化文档并重构孤立信息为结构化表格的能力。不同于传统文本转表任务依赖固定模式和狭窄领域,AOE涵盖三个多样化领域中的11个精心设计任务,要求模型根据输入查询生成上下文相关的自适应模式。实验中,我们评估了开源与闭源的前沿大模型。结果表明,即使最先进的模型也表现出显著困难。该基准已公开于 https://anonymous.4open.science/r/AOE-Benchmark/。

原文摘要 · Abstract (English)

With the emergence of large language models (LLMs), there is an expectation that LLMs can effectively extract explicit information from complex real-world documents (e.g., papers, reports). However, most LLMs generate paragraph-style answers that are chaotic, disorganized, and untraceable. To bridge this gap, we introduce the Arranged and Organized Extraction Benchmark (AOE), a new bilingual benchmark with data and documents of varying lengths designed to systematically evaluate the ability of LLMs to comprehend fragmented documents and reconstruct isolated information into one organized table. Unlike conventional text-to-table tasks, which rely on fixed schema and narrow task domains, AOE includes 11 carefully crafted tasks across three diverse domains, requiring models to generate context-specific schema tailored to varied input queries. In the experiment, we evaluated both open-source and closed-source state-of-the-art LLMs. The results show that even the most advanced models struggled significantly. The benchmark is available at https://anonymous.4open.science/r/AOE-Benchmark/.

知识提取文档理解大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。