arXiv:2511.01265cs.CL2025-11被引 2

构建首个大规模阿拉伯语金融新闻数据集,提升金融文本摘要准确性

AraFinNews: Arabic Financial Summarisation with Domain-Adapted LLMs

  • 基于10年21万组新闻标题对,构建阿拉伯语金融专有数据集
  • 领域适配模型在数字和实体信息处理上更准确,生成更连贯摘要
  • 适合研究阿拉伯语金融文本生成、大模型领域适应的学者使用

我们提出AraFinNews,目前最大公开的阿拉伯语金融新闻数据集,包含2015至2025年共212,500组新闻文章与标题对,旨在为阿拉伯语金融领域文本摘要提供真实基准。该数据集作为英文主流摘要数据集(如CNN/DailyMail)的阿拉伯语对应物,用于评估语言模型在金融领域的理解与生成能力。我们利用该数据集,研究领域特异性对抽象式摘要的影响,评测了mT5、AraT5及领域适配模型FinAraT5的表现,重点关注其在准确性、数值可靠性及专业报道风格一致性方面的能力。实验表明,领域适配模型在量化信息和实体相关表述的处理上更具优势,生成更连贯的摘要。结果强调了领域特定微调对提升阿拉伯语金融摘要叙事流畅性的重要性。数据集已开源,供非商业研究使用:https://github.com/ArabicNLP-uk/AraFinNews。

原文摘要 · Abstract (English)

We introduce AraFinNews, the largest publicly available Arabic financial news dataset to date, comprising 212,500 article-headline pairs spanning a decade of reporting from 2015 to 2025. Designed as an Arabic counterpart to major English summarisation corpora such as CNN/DailyMail, AraFinNews provides a realistic benchmark for evaluating domain-specific language understanding and generation in financial contexts. Using this resource, we investigate the impact of domain specificity on abstractive summarisation of Arabic financial texts with large language models (LLMs). In particular, we evaluate transformer-based models: mT5, AraT5, and the domain-adapted FinAraT5 to examine how financial-domain pretraining influences accuracy, numerical reliability, and stylistic alignment with professional reporting. Experimental results show that domain-adapted models generate more coherent summaries, especially in their handling of quantitative and entity-centric information. These findings highlight the importance of domain-specific adaptation for improving narrative fluency in Arabic financial summarisation. The dataset is freely available for non-commercial research at https://github.com/ArabicNLP-uk/AraFinNews.

阿拉伯语金融摘要领域适配LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。