arXiv:2412.15471cs.CL2024-12综述被引 3

梳理马拉地语自然语言处理发展脉络,盘点现有资源与技术工具。

A Review of the Marathi Natural Language Processing

  • 综述马拉地语NLP研究演进,聚焦语言特性与挑战
  • 汇总22种印度官方语言中马拉地语的最新资源与工具
  • 适合对印地语系语言处理感兴趣的科研人员参考

马拉地语是全球使用最广泛的语言之一。尽管如此,英语等主流语言的NLP进展并未迅速惠及马拉地语等印度语言。这主要源于书写系统多样性、缺乏分词策略、高质量数据集与基准测试、评估指标不足,以及马拉地语丰富的形态特征带来的挑战。自2000年代初神经网络模型兴起后,这一局面逐步改善。过去十年间,针对印度22种官方语言的语言资源建设取得显著进展。本文系统回顾了印地语系语言尤其是马拉地语的NLP研究发展,总结当前可用的先进资源与工具,并概述相关任务的技术方法。

原文摘要 · Abstract (English)

Marathi is one of the most widely used languages in the world. One might expect that the latest advances in NLP research in languages like English reach such a large community. However, NLP advancements in English didn't immediately reach Indian languages like Marathi. There were several reasons for this. They included diversity of scripts used, lack of (publicly available) resources like tokenization strategies, high quality datasets \& benchmarks, and evaluation metrics. In addition to this, the morphologically rich nature of Marathi, made NLP tasks challenging. Advances in Neural Network (NN) based models and tools since the early 2000s helped improve this situation and make NLP research more accessible. In the past 10 years, significant efforts were made to improve language resources for all 22 scheduled languages of India. This paper presents a broad overview of evolution of NLP research in Indic languages with a focus on Marathi and state-of-the-art resources and tools available to the research community. It also provides an overview of tools \& techniques associated with Marathi NLP tasks.

语言处理印度语言资源综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。