arXiv:2412.09383cs.CL2024-12中稿 · VarDial 2025被引 9

用真实语料训练序列模型,解决卢森堡语拼写变异问题

Neural Text Normalization for Luxembourgish using Real-Life Variation Data

  • 基于真实语料构建序列到序列模型,使用ByT5和mT5架构
  • 在字节、词、流水线三种模式中,序列模型表现最优
  • 适合需要定制化拼写规范的卢森堡语NLP应用

由于缺乏完整标准语体,卢森堡语文本中拼写变异现象普遍。加之标注与平行数据稀缺,且标准化进程仍在进行,开发卢森堡语NLP工具面临挑战。本文首次提出基于ByT5和mT5架构的序列到序列归一化模型,训练数据来自实际使用的词汇级变异语料。通过细致的、基于语言学的评估,测试了字节级、词级及流水线式模型在文本归一化中的优劣。结果表明,利用真实变异数据训练的序列模型能有效实现卢森堡语的定制化文本归一化。

原文摘要 · Abstract (English)

Orthographic variation is very common in Luxembourgish texts due to the absence of a fully-fledged standard variety. Additionally, developing NLP tools for Luxembourgish is a difficult task given the lack of annotated and parallel data, which is exacerbated by ongoing standardization. In this paper, we propose the first sequence-to-sequence normalization models using the ByT5 and mT5 architectures with training data obtained from word-level real-life variation data. We perform a fine-grained, linguistically-motivated evaluation to test byte-based, word-based and pipeline-based models for their strengths and weaknesses in text normalization. We show that our sequence model using real-life variation data is an effective approach for tailor-made normalization in Luxembourgish.

文本归一化卢森堡语序列模型低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。