arXiv:2409.17640cs.CLcs.AI2024-09被引 1

通过迭代训练辅助任务提升大模型长文本摘要能力。

T3: A Novel Zero-shot Transfer Learning Framework Iteratively Training on an Assistant Task for a Target Task

  • 用问答任务作为辅助,迭代优化目标摘要任务的模型。
  • 在多个数据集上实现高达14%的ROUGE提升。
  • 适合零样本迁移场景,可推广至多种任务组合。

长文本摘要对高效处理海量信息日益重要,但大型语言模型(如GPT和LLaMA系列)仍面临开源训练数据不足和上下文细节要求高的挑战。为此,我们设计了一种新型零样本迁移学习框架T3,通过在与目标任务具有结构或语义相似性的辅助任务上迭代训练基础语言模型,以提升其性能。实践中,采用问答任务作为辅助任务解决长文本摘要问题,并在BBC Summary、NarraSum、FairytaleQA和NLQuAD数据集上验证有效性,相较三种基线模型,在ROUGE上最高提升近14%,BLEU提升35%,Factscore提升16%,展现了该框架在更多辅助-目标任务组合中的潜力。

原文摘要 · Abstract (English)

Long text summarization, gradually being essential for efficiently processing large volumes of information, stays challenging for Large Language Models (LLMs) such as GPT and LLaMA families because of the insufficient open-sourced training datasets and the high requirement of contextual details dealing. To address the issue, we design a novel zero-shot transfer learning framework, abbreviated as T3, to iteratively training a baseline LLM on an assistant task for the target task, where the former should own richer data resources and share structural or semantic similarity with the latter. In practice, T3 is approached to deal with the long text summarization task by utilizing question answering as the assistant task, and further validated its effectiveness on the BBC summary, NarraSum, FairytaleQA, and NLQuAD datasets, with up to nearly 14% improvement in ROUGE, 35% improvement in BLEU, and 16% improvement in Factscore compared to three baseline LLMs, demonstrating its potential for more assistant-target task combinations.

零样本学习摘要生成迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。