为高棉语设计的轻量级生成模型,提升文本自然度和生成质量。
PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation
- 从头训练,融合分词与标准化等语言模块提升高棉语处理能力
- 在机器翻译、摘要生成等任务上优于mBART50模型
- 特别优化了空格使用,更符合高棉语书写习惯,适合本地化应用
本文提出PrahokBART,一个专为高棉语设计的轻量级预训练序列到序列模型,基于精心筛选的高棉语与英语语料库从零开始训练。通过引入分词、标准化等语言组件,有效解决现有多语言模型忽略的高棉语语言问题。在机器翻译、文本摘要和标题生成三项生成任务上,PrahokBART表现优于mBART50。分析还揭示了各语言模块的影响,并评估了模型在生成过程中对空格的处理能力,这对高棉语文本自然性至关重要。
原文摘要 · Abstract (English)
This work introduces {\it PrahokBART}, a compact pre-trained sequence-to-sequence model trained from scratch for Khmer using carefully curated Khmer and English corpora. We focus on improving the pre-training corpus quality and addressing the linguistic issues of Khmer, which are ignored in existing multilingual models, by incorporating linguistic components such as word segmentation and normalization. We evaluate PrahokBART on three generative tasks: machine translation, text summarization, and headline generation, where our results demonstrate that it outperforms mBART50, a strong multilingual pre-trained model. Additionally, our analysis provides insights into the impact of each linguistic module and evaluates how effectively our model handles space during text generation, which is crucial for the naturalness of texts in Khmer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。