arXiv:2410.24140cs.CL2024-10被引 10

忽略变音符号会损害多语言模型性能,应统一处理以提升公平性。

Don't Touch My Diacritics

  • 提出统一处理文本中的变音符号,避免预处理导致偏差。
  • 案例表明,随意移除或编码不一致会降低模型表现。
  • 适合关注多语言NLP公平性的研究者与开发者参考。

在将文本输入NLP模型前的预处理环节中,存在诸多决策点,可能对模型性能产生意外影响。本文聚焦于多种语言和文字系统中变音符号的处理问题。通过多个案例研究,我们展示了变音字符编码不一致以及完全去除变音符号所带来的负面影响。呼吁社区在所有模型和工具包中采取简单但必要的措施,以改进对带变音符号文本的处理,从而提升多语言NLP的公平性。

原文摘要 · Abstract (English)

The common practice of preprocessing text before feeding it into NLP models introduces many decision points which have unintended consequences on model performance. In this opinion piece, we focus on the handling of diacritics in texts originating in many languages and scripts. We demonstrate, through several case studies, the adverse effects of inconsistent encoding of diacritized characters and of removing diacritics altogether. We call on the community to adopt simple but necessary steps across all models and toolkits in order to improve handling of diacritized text and, by extension, increase equity in multilingual NLP.

多语言NLP变音符号模型公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。