用图算法给孟加拉语词汇排序,提升文本摘要与信息检索效果。
Ranking of Bangla Word Graph using Graph-based Ranking Algorithms
- 构建孟加拉语文本的词汇图,用图算法计算词重要性。
- 在真实数据上测试,各算法F1值表现各异,可比较优劣。
- 适用于语言处理研究者,尤其关注低资源语言的场景。
词语排序是文本摘要和信息检索的重要手段。词图将句子或文本中的词语表示为图的顶点,并展示词语间的关系,有助于判断词在图中的相对重要性。本研究利用多种基于图的排名算法,对孟加拉语文本中的词语进行排序。由于缺乏标准的孟加拉语词库,研究采用印度语言词性标注语料库(Indian Language POS-tag Corpora),该语料库包含大量带词性标注的孟加拉语句子。为应用词图于各类图算法,研究执行了标准化预处理流程,包括文本清洗、分词与词性标注等步骤。随后将处理后的词图输入不同图排名算法,进行性能对比。实验基于真实数据,通过F1分数评估各算法的准确性,全面展示了词图构建与排名的完整流程。
原文摘要 · Abstract (English)
Ranking words is an important way to summarize a text or to retrieve information. A word graph is a way to represent the words of a sentence or a text as the vertices of a graph and to show the relationship among the words. It is also useful to determine the relative importance of a word among the words in the word-graph. In this research, the ranking of Bangla words are calculated, representing Bangla words from a text in a word graph using various graph based ranking algorithms. There is a lack of a standard Bangla word database. In this research, the Indian Language POS-tag Corpora is used, which has a rich collection of Bangla words in the form of sentences with their parts of speech tags. For applying a word graph to various graph based ranking algorithms, several standard procedures are applied. The preprocessing steps are done in every word graph and then applied to graph based ranking algorithms to make a comparison among these algorithms. This paper illustrate the entire procedure of calculating the ranking of Bangla words, including the construction of the word graph from text. Experimental result analysis on real data reveals the accuracy of each ranking algorithm in terms of F1 measure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。