arXiv:2506.11602cs.CLcs.AI2025-06中稿 · LREC 2026被引 2

LLMs在阿拉伯语和约鲁巴语文本加音标上表现优异,但小模型易产生幻觉。

Are LLMs Good Text Diacritizers? An Arabic and Yoruba Case Study

  • 使用多语言数据集MultiDiac评估12个LLM和4个专用模型的加音标能力。
  • 多数现成LLM性能超过专用模型,但小型模型存在明显幻觉问题。
  • 小模型经LoRA微调后,约鲁巴语加音标效果提升且幻觉减少。

我们研究了大语言模型(LLMs)在两种语言类型差异显著的语言——阿拉伯语和约鲁巴语中的文本加音标效果。为实现严谨评估,我们引入了一个新的多语言数据集MultiDiac,包含多样样本,涵盖各类音标歧义情况。我们评估了12个规模、可访问性和语言覆盖范围各异的LLM,并与4个专用加音标模型进行对比。此外,我们还使用LoRA对4个小型开源模型在约鲁巴语上进行了微调。结果表明,许多现成的LLM表现优于专用模型,但小型模型容易出现幻觉。在小数据集上微调可有效提升约鲁巴语加音标性能并减少幻觉。

原文摘要 · Abstract (English)

We investigate the effectiveness of large language models (LLMs) for text diacritization in two typologically distinct languages: Arabic and Yoruba. To enable a rigorous evaluation, we introduce a novel multilingual dataset MultiDiac, with diverse samples that capture a range of diacritic ambiguities. We evaluate 12 LLMs varying in size, accessibility, and language coverage, and benchmark them against $4$ specialized diacritization models. Additionally, we fine-tune four small open-source models using LoRA for Yoruba. Our results show that many off-the-shelf LLMs outperform specialized diacritization models, but smaller models suffer from hallucinations. We find that fine-tuning on a small dataset can help improve diacritization performance and reduce hallucinations for Yoruba.

文本加音标大语言模型约鲁巴语阿拉伯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。