测试不同摘要模型在跨领域时的失效原因,发现事实性下降是主要问题。
Can one size fit all?: Measuring Failure in Multi-Document Summarization Domain Transfer
- 对比四种训练方法在跨领域时的表现差异
- 发现科学和对话领域摘要事实性下降超过30%
- 提醒直接使用评测指标可能产生误导
抽象型多文档摘要(MDS)旨在从新闻、对话等多篇文档中自动提炼信息。当前主流训练方法包括:端到端预训练(直接法)、分块再摘要、提取再摘要,以及基于GPT风格的推理方法。本文评估了不同模型在新闻、科学和对话三类领域间的零样本跨域迁移表现,分析其失败原因。定义跨域失败为事实性降低、与参考摘要偏差增大及整体质量下降。结果表明,模型在非训练领域上事实性平均下降超30%,尤其在科学与对话领域更显著。同时指出,现有主流摘要评测指标若不加调整直接应用,可能导致对模型性能的误判。
原文摘要 · Abstract (English)
Abstractive multi-document summarization (MDS) is the task of automatically summarizing information in multiple documents, from news articles to conversations with multiple speakers. The training approaches for current MDS models can be grouped into four approaches: end-to-end with special pre-training ("direct"), chunk-then-summarize, extract-then-summarize, and inference with GPT-style models. In this work, we evaluate MDS models across training approaches, domains, and dimensions (reference similarity, quality, and factuality), to analyze how and why models trained on one domain can fail to summarize documents from another (News, Science, and Conversation) in the zero-shot domain transfer setting. We define domain-transfer "failure" as a decrease in factuality, higher deviation from the target, and a general decrease in summary quality. In addition to exploring domain transfer for MDS models, we examine potential issues with applying popular summarization metrics out-of-the-box.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。