用LSTM等模型实现锡尔赫特语到现代孟加拉语的精准翻译。
Improving Bangla Linguistics: Advanced LSTM, Bi-LSTM, and Seq2Seq Models for Translating Sylheti to Modern Bangla
- 构建NLP系统,用三种模型将现代孟加拉语转为锡尔赫特方言。
- 1200条数据训练,LSTM模型准确率达89.3%表现最佳。
- 助力地方语言数字化,适合孟加拉语NLP研究者参考。
孟加拉语是孟加拉国的国语,但各地口音差异大,如锡尔赫特地区使用当地方言锡尔赫特语。尽管近年来有部分关于孟加拉语的自然语言处理研究(如情感分析、假新闻检测),但针对地方语言的研究仍较少。本研究聚焦于锡尔赫特语,提出一个综合NLP系统,实现从纯正或现代孟加拉语到本地锡尔赫特语的翻译。使用1200条数据训练了LSTM、Bi-LSTM和Seq2Seq三种模型,其中LSTM表现最优,准确率达到89.3%。研究成果可为未来孟加拉语NLP的发展提供基础支持。
原文摘要 · Abstract (English)
Bangla or Bengali is the national language of Bangladesh, people from different regions don't talk in proper Bangla. Every division of Bangladesh has its own local language like Sylheti, Chittagong etc. In recent years some papers were published on Bangla language like sentiment analysis, fake news detection and classifications, but a few of them were on Bangla languages. This research is for the local language and this particular paper is on Sylheti language. It presented a comprehensive system using Natural Language Processing or NLP techniques for translating Pure or Modern Bangla to locally spoken Sylheti Bangla language. Total 1200 data used for training 3 models LSTM, Bi-LSTM and Seq2Seq and LSTM scored the best in performance with 89.3% accuracy. The findings of this research may contribute to the growth of Bangla NLP researchers for future more advanced innovations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。