将复杂文本分解为原子语句,自动提取结构化表格。
Map&Make: Schema Guided Text to Table Generation
- 通过分解文本为原子命题,挖掘隐藏的表结构
- 在Rotowire和Livesum上显著提升生成质量
- 适合需要精准结构化摘要的科研与业务场景
将密集、详细的非结构化文本转化为可解释的摘要表格,即文本到表格生成,是信息检索中的关键任务。现有方法往往忽视复杂信息的提取方式与内容选择,且缺乏从文本中推断数据的能力。本文提出一种通用方法Map&Make,将文本“解构”为命题级原子语句,实现细粒度分解以提取潜在表结构。该结构用于填充表格,准确捕捉原始文本中的定性细节与定量事实。我们在两个具有挑战性的数据集Rotowire(以复杂多表结构著称)和Livesum(需数值聚合)上测试该方法。通过仔细识别并修正Rotowire中的幻觉错误,我们构建了更清洁可靠的基准。采用全面的对比与无参考评估指标进行严谨评测。结果表明,在两个数据集上均有显著性能提升,且生成结果更具可解释性。通过详尽的消融实验与分析,我们揭示了性能优势的关键因素,并验证了该框架在结构化摘要任务中的实用性。
原文摘要 · Abstract (English)
Transforming dense, detailed, unstructured text into an interpretable and summarised table, also colloquially known as Text-to-Table generation, is an essential task for information retrieval. Current methods, however, miss out on how and what complex information to extract; they also lack the ability to infer data from the text. In this paper, we introduce a versatile approach, Map&Make, which "dissects" text into propositional atomic statements. This facilitates granular decomposition to extract the latent schema. The schema is then used to populate the tables that capture the qualitative nuances and the quantitative facts in the original text. Our approach is tested against two challenging datasets, Rotowire, renowned for its complex and multi-table schema, and Livesum, which demands numerical aggregation. By carefully identifying and correcting hallucination errors in Rotowire, we aim to achieve a cleaner and more reliable benchmark. We evaluate our method rigorously on a comprehensive suite of comparative and referenceless metrics. Our findings demonstrate significant improvement results across both datasets with better interpretability in Text-to-Table generation. Moreover, through detailed ablation studies and analyses, we investigate the factors contributing to superior performance and validate the practicality of our framework in structured summarization tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。