39种语言的跨语言摘要实验表明,分步翻译再摘要优于端到端模型。
With Good MT There is No Need For End-to-End: A Case for Translate-then-Summarize Cross-lingual Summarization
- 采用先翻译后摘要的流水线设计,简单有效。
- 在39种语言上均优于使用海量平行数据的端到端模型。
- 可通过BLEU分数预判语言对是否适合该方法,指导实践。
近期研究认为,跨语言摘要的端到端系统在性能上可与传统流水线设计比肩甚至更优。然而深入分析发现,这一结论仅基于少数语言或低性能的流水线基线。本文在39种源语言转英语的跨语言摘要任务上对比了两种范式,结果表明,一个简单的‘翻译-再摘要’流水线设计始终优于即使拥有大量平行数据的端到端系统。对于表现不佳的语言,我们发现系统性能与公开发布的BLEU分数高度相关,使从业者可事先评估语言对的可行性。与当前研究趋势相反,我们的结果表明,单语摘要和翻译任务的独立进步,整体效果优于端到端系统,提示应谨慎设计端到端方案。
原文摘要 · Abstract (English)
Recent work has suggested that end-to-end system designs for cross-lingual summarization are competitive solutions that perform on par or even better than traditional pipelined designs. A closer look at the evidence reveals that this intuition is based on the results of only a handful of languages or using underpowered pipeline baselines. In this work, we compare these two paradigms for cross-lingual summarization on 39 source languages into English and show that a simple \textit{translate-then-summarize} pipeline design consistently outperforms even an end-to-end system with access to enormous amounts of parallel data. For languages where our pipeline model does not perform well, we show that system performance is highly correlated with publicly distributed BLEU scores, allowing practitioners to establish the feasibility of a language pair a priori. Contrary to recent publication trends, our result suggests that the combination of individual progress of monolingual summarization and translation tasks offers better performance than an end-to-end system, suggesting that end-to-end designs should be considered with care.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。