综述阿拉伯语混用语言技术的研究进展与挑战
A Survey of Code-switched Arabic NLP: Progress, Challenges, and Future Directions
- 系统梳理阿拉伯语混用文本的自然语言处理研究
- 指出多语言混合导致模型泛化困难的核心挑战
- 适合关注中东语言技术与跨语言研究的学者
阿拉伯世界语言环境呈现复杂的双语与多语特征,包括现代标准阿拉伯语、多种方言及次方言,以及多种欧洲语言。这一多样化的语言格局催生了阿拉伯语内部及阿拉伯语与外语之间的语言混用现象。该现象在区域内广泛存在,因此在开发语言技术时必须应对这种语言需求。本文综述了当前阿拉伯语混用自然语言处理领域的文献,从整体视角出发,呈现当前研究进展、面临挑战、研究空白,并提出未来研究方向的建议。
原文摘要 · Abstract (English)
Language in the Arab world presents a complex diglossic and multilingual setting, involving the use of Modern Standard Arabic, various dialects and sub-dialects, as well as multiple European languages. This diverse linguistic landscape has given rise to code-switching, both within Arabic varieties and between Arabic and foreign languages. The widespread occurrence of code-switching across the region makes it vital to address these linguistic needs when developing language technologies. In this paper, we provide a review of the current literature in the field of code-switched Arabic NLP, offering a broad perspective on ongoing efforts, challenges, research gaps, and recommendations for future research directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。