arXiv:2605.05392cs.CLcs.AI2026-05

从无查询摘要数据自动生成精准查询,提升问答式摘要效果

Generating Query-Focused Summarization Datasets from Query-Free Summarization Datasets

论文配图:Generating Query-Focused Summarization Datasets from Query-Free Summarization Datasets
图 1 · 摘自论文原文
  • 基于证据生成模型自动提取查询关键词,无需人工标注
  • 生成查询使摘要模型在ROUGE指标上达到与原始查询相当的水平
  • 适合需要构建查询聚焦摘要数据集的研究者和应用开发者

大规模摘要数据集广泛用于摘要任务,但通常不包含查询。针对查询聚焦摘要(QFS)任务,本文提出两个核心问题:能否从无查询数据集中自动生成基于证据的查询关键词?基于证据的查询生成是否有助于提升QFS性能?为此,本文提出一种基于证据的查询生成模型。通过对比两个QFS数据集中原始查询与系统生成查询的相似性,评估模型内在质量;同时使用多种预训练模型及SOTA QFS模型进行摘要实验,验证生成查询的外在表现。实验表明,采用证据生成查询所生成的摘要,在ROUGE得分上与使用原始查询的结果具有竞争力。

原文摘要 · Abstract (English)

Large-scale datasets are widely used to perform summarization tasks, but they may not include queries alongside documents and summaries. In the search for suitable datasets for Query-Focused Summarization (QFS), we identify two research questions: Is it possible to automatically generate evidence-based query keywords from query-free datasets? Does evidence-based query generation support the QFS task? This paper proposes an evidence-based model to generate queries from query-free datasets. To evaluate our model intrinsically, we compare the similarity between the original queries and the system-generated queries of two QFS datasets. We also perform summarization tasks using different pre-trained models, as well as a state-of-the-art (SOTA) QFS model, to measure the extrinsic performance of our query generation approach. Experimental results indicate that summaries generated using evidence-based queries achieve competitive ROUGE scores compared to those generated from the original queries.

摘要生成查询生成数据构建ROUGE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。