用字节级模型统一处理梵语多种自然语言任务,性能超越传统方法。
One Model is All You Need: ByT5-Sanskrit, a Unified Model for Sanskrit NLP Tasks
- 基于字节的预训练模型,无需依赖外部词典资源。
- 在梵语分词、依存句法分析等任务上达到新最优,提升2-5个百分点。
- 适用于梵语标注、信息检索及机器翻译前处理,通用性强。
形态丰富的语言在下游自然语言处理中极具挑战性。本文提出针对梵语的预训练语言模型 ByT5-Sanskrit,用于处理该形态复杂语言的 NLP 任务。在已有的梵语分词任务上,其性能显著优于以往数据驱动方法,并达到当前最佳词典基模型水平。部署更简便,对未覆盖于外部语言资源的数据更具鲁棒性。此外,在吠陀梵语依存句法分析与光学字符识别后纠错任务中也取得新最佳结果。基于数字梵语语料库,我们构建了一个新型多任务数据集,用于联合训练梵语分词、词形还原与语法标注任务。在此数据集上微调 ByT5-Sanskrit,形成可应用于多种下游任务的多功能模型。该模型已成功应用于梵语语言学标注项目、信息检索系统以及梵语机器翻译流程中的预处理环节。实验还表明,该方法在其他形态丰富语言的词形还原与依存句法分析任务上同样获得最佳成绩。研究证明,字节级预训练模型在形态丰富语言上表现优异,优于基于分词器的模型,为构建此类语言的 NLP 系统提供了重要方向。
原文摘要 · Abstract (English)
Morphologically rich languages are notoriously challenging to process for downstream NLP applications. This paper presents a new pretrained language model, ByT5-Sanskrit, designed for NLP applications involving the morphologically rich language Sanskrit. We evaluate ByT5-Sanskrit on established Sanskrit word segmentation tasks, where it outperforms previous data-driven approaches by a considerable margin and matches the performance of the current best lexicon-based model. It is easier to deploy and more robust to data not covered by external linguistic resources. It also achieves new state-of-the-art results in Vedic Sanskrit dependency parsing and OCR post-correction tasks. Additionally, based on the Digital Corpus of Sanskrit, we introduce a novel multitask dataset for the joint training of Sanskrit word segmentation, lemmatization, and morphosyntactic tagging tasks. We fine-tune ByT5-Sanskrit on this dataset, creating a versatile multitask model for various downstream Sanskrit applications. We have used this model in Sanskrit linguistic annotation projects, in information retrieval setups, and as a preprocessing step in a Sanskrit machine translation pipeline. We also show that our approach yields new best scores for lemmatization and dependency parsing of other morphologically rich languages. We thus demonstrate that byte-level pretrained language models can achieve excellent performance for morphologically rich languages, outperforming tokenizer-based models and presenting an important vector of exploration when constructing NLP pipelines for such languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。