arXiv:2412.20223cs.CL2024-12中稿 · ICLR被引 2

构建非洲语言新闻标题生成数据集,验证本地模型优于通用大模型。

AfriHG: News headline generation for African Languages

  • 整合XLSum与MasakhaNEWS数据,构建16种非洲语言的新闻标题数据集。
  • 非洲专用模型AfriTeVa V2在16种语言上表现优于mT5-base模型。
  • 313M参数的微调模型性能媲美130亿参数的大模型提示调用。

本文提出AfriHG——一个结合XLSum和MasakhaNEWS数据集构建的新闻标题生成数据集,覆盖非洲广泛使用的16种语言。我们测试了两种序列到序列模型(mT5-base和AfriTeVa V2)以及Aya-101大语言模型。实验结果表明,以非洲为中心的序列到序列模型如AfriTeVa V2在多数语言上优于大规模多语言模型mT5-base。最终结果显示,使用313M参数的AfriTeVa V2进行微调后的性能,可与使用超过130亿参数的Aya-101大模型通过提示工程实现的性能相媲美。

原文摘要 · Abstract (English)

This paper introduces AfriHG -- a news headline generation dataset created by combining from XLSum and MasakhaNEWS datasets focusing on 16 languages widely spoken by Africa. We experimented with two seq2eq models (mT5-base and AfriTeVa V2), and Aya-101 LLM. Our results show that Africa-centric seq2seq models such as AfriTeVa V2 outperform the massively multilingual mT5-base model. Finally, we show that the performance of fine-tuning AfriTeVa V2 with 313M parameters is competitive to prompting Aya-101 LLM with more than 13B parameters.

新闻生成非洲语言小模型多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。