用大模型提升捷克历史文献摘要,填补数据与技术空白
Large Language Models for Summarizing Czech Historical Documents and Beyond
- 使用Mistral和mT5模型处理捷克语摘要任务
- 在SumeCzech数据集上达到新最好结果
- 发布首个历史捷克文献摘要数据集,适合语言研究者
文本摘要旨在将长篇文本浓缩为简明版本,同时保留核心意义与关键信息。尽管英语及其他高资源语言的摘要研究已取得显著进展,但捷克语,特别是历史文献的摘要仍因语言复杂性和标注数据稀缺而研究不足。大型语言模型如Mistral和mT5已在多种自然语言处理任务和语言中表现优异。本文采用这些模型进行捷克语摘要,取得两项主要贡献:(1)在现代捷克语摘要数据集SumeCzech上实现新的最优性能;(2)提出一个全新的历史捷克文献摘要数据集Posel od Čerchova,并提供基线结果。这两项成果为推进捷克语文本摘要奠定了基础,也为捷克历史文本处理开辟了新方向。
原文摘要 · Abstract (English)
Text summarization is the task of shortening a larger body of text into a concise version while retaining its essential meaning and key information. While summarization has been significantly explored in English and other high-resource languages, Czech text summarization, particularly for historical documents, remains underexplored due to linguistic complexities and a scarcity of annotated datasets. Large language models such as Mistral and mT5 have demonstrated excellent results on many natural language processing tasks and languages. Therefore, we employ these models for Czech summarization, resulting in two key contributions: (1) achieving new state-of-the-art results on the modern Czech summarization dataset SumeCzech using these advanced models, and (2) introducing a novel dataset called Posel od Čerchova for summarization of historical Czech documents with baseline results. Together, these contributions provide a great potential for advancing Czech text summarization and open new avenues for research in Czech historical text processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。