arXiv:2409.12134cs.CLcs.AI2024-09被引 2

融合提取与生成的越南语多文档摘要框架,效果优于现有方法。

BERT-VBD: Vietnamese Multi-Document Summarization Framework

  • 双阶段流程:先用改进BERT提取关键句,再用VBD-LLaMA2生成摘要
  • 在VN-MDS数据集上达到ROUGE-2 39.6%,超越当前最佳模型
  • 专为越南语设计,适合多文档摘要和低资源语言研究者

针对多文档摘要(MDS)挑战,已有众多提取式与生成式方法被提出,但各有局限,单一依赖任一方式效果有限。一种新兴且有前景的策略是融合提取与生成方法。尽管该领域研究丰富,但在越南语处理中的联合方法仍较少。本文提出一种新型越南语多文档摘要框架,采用两组件流水线架构,整合提取与生成技术。第一组件通过改进的预训练BERT网络,利用孪生与三元组网络结构生成语义有意义的短语嵌入,识别每篇文档的关键句子。第二组件采用VBD-LLaMA2-7B-50b模型进行抽象生成,最终形成摘要文档。所提框架表现良好,在VN-MDS数据集上取得ROUGE-2 39.6%的分数,优于现有最先进基线。

原文摘要 · Abstract (English)

In tackling the challenge of Multi-Document Summarization (MDS), numerous methods have been proposed, spanning both extractive and abstractive summarization techniques. However, each approach has its own limitations, making it less effective to rely solely on either one. An emerging and promising strategy involves a synergistic fusion of extractive and abstractive summarization methods. Despite the plethora of studies in this domain, research on the combined methodology remains scarce, particularly in the context of Vietnamese language processing. This paper presents a novel Vietnamese MDS framework leveraging a two-component pipeline architecture that integrates extractive and abstractive techniques. The first component employs an extractive approach to identify key sentences within each document. This is achieved by a modification of the pre-trained BERT network, which derives semantically meaningful phrase embeddings using siamese and triplet network structures. The second component utilizes the VBD-LLaMA2-7B-50b model for abstractive summarization, ultimately generating the final summary document. Our proposed framework demonstrates a positive performance, attaining ROUGE-2 scores of 39.6% on the VN-MDS dataset and outperforming the state-of-the-art baselines.

多文档摘要越南语生成式摘要BERT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。