构建阿拉伯语专有名词带调音标注数据集,助力外文名词发音还原。
Proper Noun Diacritization for Arabic Wikipedia: A Benchmark Dataset
- 人工标注阿拉伯语专有名词的完整调音符号,附英文对应词。
- GPT-4o在还原任务中准确率达73%,显示挑战仍大。
- 适合研究阿拉伯语命名实体、语音还原与跨语言对齐的学者。
阿拉伯语维基百科中的专有名词常缺少调音符号,导致发音和理解歧义,尤其针对外来语源的转写词。尽管阿拉伯语的转写与调音已分别得到充分研究,但二者交叉问题仍被忽视。本文构建了一个手动标注的阿拉伯语专有名词调音数据集,涵盖多种来源,并提供其对应的英文维基百科释义。我们介绍了构建过程中的挑战与规范,并以GPT-4o为基准,在给定无调音阿拉伯语和英文形式的前提下,测试其恢复完整调音的能力。实验结果显示73%的准确率,表明该任务难度高,亟需更优模型与资源。数据集已公开,以推动相关研究。
原文摘要 · Abstract (English)
Proper nouns in Arabic Wikipedia are frequently undiacritized, creating ambiguity in pronunciation and interpretation, especially for transliterated named entities of foreign origin. While transliteration and diacritization have been well-studied separately in Arabic NLP, their intersection remains underexplored. In this paper, we introduce a new manually diacritized dataset of Arabic proper nouns of various origins with their English Wikipedia equivalent glosses, and present the challenges and guidelines we followed to create it. We benchmark GPT-4o on the task of recovering full diacritization given the undiacritized Arabic and English forms, and analyze its performance. Achieving 73% accuracy, our results underscore both the difficulty of the task and the need for improved models and resources. We release our dataset to facilitate further research on Arabic Wikipedia proper noun diacritization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。