针对孟加拉语资源匮乏问题,系统梳理其词干提取研究现状与挑战
Stemming -- The Evolution and Current State with a Focus on Bangla
- 综述孟加拉语词干提取方法,聚焦形态变体处理
- 指出现有研究缺乏可复现实现和有效评估指标
- 呼吁开发鲁棒词干提取器以支持低资源语言处理
孟加拉语是全球第七大使用语言,拥有3亿母语者,但因资源有限和标注数据集缺失,在数字领域存在代表性不足。词干提取作为语言分析的关键预处理步骤,对高度屈折的低资源语言如孟加拉语至关重要,能显著减少算法需处理的词数。本文全面综述了孟加拉语词干提取方法,强调有效处理形态变体的重要性。研究发现,现有文献存在明显断层,且缺乏可访问的实现代码供复现。同时,评估方法也存在缺陷,亟需更相关的评价指标。鉴于孟加拉语丰富的形态结构和方言多样性,本文指出了其带来的挑战,并提出未来词干提取器发展的方向。最后,呼吁开发更稳健的孟加拉语词干提取工具,推动该领域的持续研究,提升语言分析能力。
原文摘要 · Abstract (English)
Bangla, the seventh most widely spoken language worldwide with 300 million native speakers, faces digital under-representation due to limited resources and lack of annotated datasets. Stemming, a critical preprocessing step in language analysis, is essential for low-resource, highly-inflectional languages like Bangla, because it can reduce the complexity of algorithms and models by significantly reducing the number of words the algorithm needs to consider. This paper conducts a comprehensive survey of stemming approaches, emphasizing the importance of handling morphological variants effectively. While exploring the landscape of Bangla stemming, it becomes evident that there is a significant gap in the existing literature. The paper highlights the discontinuity from previous research and the scarcity of accessible implementations for replication. Furthermore, it critiques the evaluation methodologies, stressing the need for more relevant metrics. In the context of Bangla's rich morphology and diverse dialects, the paper acknowledges the challenges it poses. To address these challenges, the paper suggests directions for Bangla stemmer development. It concludes by advocating for robust Bangla stemmers and continued research in the field to enhance language analysis and processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。