arXiv:2511.18848cs.CL2025-11

用大模型提升捷克语文本摘要,覆盖现代与历史文献

Large Language Models for the Summarization of Czech Documents: From History to the Present

  • 用Mistral和mT5等大模型处理捷克语摘要,适配语法复杂的中资源语言
  • 在SumeCzech数据集上达到新最好效果,历史文本数据集达19世纪出版物
  • 开源历史文本摘要数据集Posel od Čerchova,助力低资源语言研究

文本摘要旨在自动将长文本压缩为简洁连贯的摘要,同时保留原意和关键信息。尽管该任务在英语等高资源语言中已有广泛研究,但捷克语摘要,特别是历史文档摘要,仍研究不足,主要因捷克语语言复杂且缺乏高质量标注数据。本文利用大型语言模型(LLMs),如Mistral和mT5,其在多语言任务中表现优异。我们还提出一种基于翻译的方法:先将捷克文译为英文,用英文模型生成摘要,再回译为捷克语。主要贡献包括:在现代捷克语摘要基准SumeCzech上实现新最佳性能,证明多语言大模型对形态丰富、中资源语言的有效性;构建了新数据集Posel od Čerchova,源自19世纪数字化报刊,用于历史文本抽象摘要;提供现代大模型基线,推动该领域的进一步研究。本工作为捷克语摘要及低资源语言摘要研究奠定基础。

原文摘要 · Abstract (English)

Text summarization is the task of automatically condensing longer texts into shorter, coherent summaries while preserving the original meaning and key information. Although this task has been extensively studied in English and other high-resource languages, Czech summarization, particularly in the context of historical documents, remains underexplored. This is largely due to the inherent linguistic complexity of Czech and the lack of high-quality annotated datasets. In this work, we address this gap by leveraging the capabilities of Large Language Models (LLMs), specifically Mistral and mT5, which have demonstrated strong performance across a wide range of natural language processing tasks and multilingual settings. In addition, we also propose a translation-based approach that first translates Czech texts into English, summarizes them using an English-language model, and then translates the summaries back into Czech. Our study makes the following main contributions: We demonstrate that LLMs achieve new state-of-the-art results on the SumeCzech dataset, a benchmark for modern Czech text summarization, showing the effectiveness of multilingual LLMs even for morphologically rich, medium-resource languages like Czech. We introduce a new dataset, Posel od Čerchova, designed for the summarization of historical Czech texts. This dataset is derived from digitized 19th-century publications and annotated for abstractive summarization. We provide initial baselines using modern LLMs to facilitate further research in this underrepresented area. By combining cutting-edge models with both modern and historical Czech datasets, our work lays the foundation for further progress in Czech summarization and contributes valuable resources for future research in Czech historical document processing and low-resource summarization more broadly.

文本摘要大模型历史文献捷克语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。