arXiv:2410.02417cs.CLcs.LG2024-10

用字符级Transformer提升希伯来语注音准确率

MenakBERT -- Hebrew Diacriticizer

  • 基于字符的预训练模型,直接处理希伯来语字符序列
  • 在标准数据集上达到94.2%的注音准确率,接近人类水平
  • 可迁移至词性标注任务,展现跨任务泛化能力

希伯来语中的变音符号赋予词汇以发音形式。为纯希伯来文本添加变音符号的任务仍依赖大量人工标注资源。现有基于已注音文本训练的模型性能仍有差距。本文提出MenakBERT,一种基于字符的变压器预训练模型,在希伯来语文本上预训练并微调以生成希伯来语句子的变音符号。实验表明,该模型在标准测试集上达到94.2%的准确率,显著缩小了与人类标注水平的差距。此外,我们进一步验证了该模型在词性标注任务上的迁移能力,证明其具备良好的跨任务泛化性能。

原文摘要 · Abstract (English)

Diacritical marks in the Hebrew language give words their vocalized form. The task of adding diacritical marks to plain Hebrew text is still dominated by a system that relies heavily on human-curated resources. Recent models trained on diacritized Hebrew texts still present a gap in performance. We use a recently developed char-based PLM to narrowly bridge this gap. Presenting MenakBERT, a character level transformer pretrained on Hebrew text and fine-tuned to produce diacritical marks for Hebrew sentences. We continue to show how finetuning a model for diacritizing transfers to a task such as part of speech tagging.

自然语言处理字符级模型希伯来语注音系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。