arXiv:2505.03481cs.CL2025-05被引 1

用句子嵌入提升长文本摘要质量,效果优于传统方法。

Sentence Embeddings as an intermediate target in end-to-end summarisation

  • 将句子级嵌入作为中间目标,结合抽取与抽象模型
  • 在大规模住宿评价数据上超越现有方法,显著提升摘要质量
  • 适合处理源文本与目标摘要对齐不严的场景

当前基于神经网络的文档摘要方法在处理大规模输入数据集时表现不佳。本文针对住宿评价的端到端摘要任务,提出一种新方法:通过结合抽取式策略与外部预训练的句子级嵌入,并融入抽象摘要模型,有效提升了摘要质量。实验表明,在大规模输入数据上,该方法优于现有方法。此外,我们证明以预测摘要的句子级嵌入作为目标,相比传统预测句子选择概率分布的方式,能显著提升松散对齐的源-目标语料上的端到端系统性能。

原文摘要 · Abstract (English)

Current neural network-based methods to the problem of document summarisation struggle when applied to datasets containing large inputs. In this paper we propose a new approach to the challenge of content-selection when dealing with end-to-end summarisation of user reviews of accommodations. We show that by combining an extractive approach with externally pre-trained sentence level embeddings in an addition to an abstractive summarisation model we can outperform existing methods when this is applied to the task of summarising a large input dataset. We also prove that predicting sentence level embedding of a summary increases the quality of an end-to-end system for loosely aligned source to target corpora, than compared to commonly predicting probability distributions of sentence selection.

摘要生成句子嵌入端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。