arXiv:2412.20218cs.CL2024-12中稿 · ICLR被引 3
用T5提升约鲁巴语自动加声调,数据与模型越大效果越好
YAD: Leveraging T5 for Improved Automatic Diacritization of Yorùbá Text
- 用约鲁巴语微调T5模型,实现端到端声调预测
- 新数据集上性能超越多语言T5模型,准确率超90%
- 适合低资源语言处理、非洲语言计算的研究者
本文提出约鲁巴语自动加声调(YAD)基准数据集,用于评估相关系统。我们针对约鲁巴语预训练了文本到文本的T5模型,并证明该模型优于多个多语言T5模型。此外,实验表明更多数据和更大模型在约鲁巴语加声调任务中表现更优。
原文摘要 · Abstract (English)
In this work, we present Yorùbá automatic diacritization (YAD) benchmark dataset for evaluating Yorùbá diacritization systems. In addition, we pre-train text-to-text transformer, T5 model for Yorùbá and showed that this model outperform several multilingually trained T5 models. Lastly, we showed that more data and larger models are better at diacritization for Yorùbá
自然语言处理低资源语言T5声调标注
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。