首次为低资源语言纳加语构建情感分析系统
Sentiment Analysis and Emotion Classification using Machine Learning Techniques for Nagamese Language -- A Low-resource Language
- 基于1195个纳加语词汇构建情感词典,结合机器学习特征
- 在纳加语文本上实现正/负/中性情感分类,首次尝试
- 适合低资源语言自然语言处理研究者参考
纳加语(Nagamese),又称那加皮钦语,是一种以阿萨姆语为词源的克里奥尔语,主要作为印度东北部那加兰邦与阿萨姆邦人民之间贸易交流的沟通工具。已有大量情感分析研究集中于英语、印地语等资源丰富语言,但对纳加语尚无相关工作。据我们所知,本文是首个针对纳加语进行情感分析与情绪分类的研究。研究旨在识别纳加语文本中的情感极性(正面、负面、中性)及基本情绪。我们构建了一个包含1,195个纳加语词汇的情感极性词典,并结合其他特征,使用朴素贝叶斯和支持向量机等监督学习方法进行建模。该研究填补了低资源语言情感分析领域的空白。
原文摘要 · Abstract (English)
The Nagamese language, a.k.a Naga Pidgin, is an Assamese-lexified creole language developed primarily as a means of communication in trade between the people from Nagaland and people from Assam in the north-east India. Substantial amount of work in sentiment analysis has been done for resource-rich languages like English, Hindi, etc. However, no work has been done in Nagamese language. To the best of our knowledge, this is the first attempt on sentiment analysis and emotion classification for the Nagamese Language. The aim of this work is to detect sentiments in terms of polarity (positive, negative and neutral) and basic emotions contained in textual content of Nagamese language. We build sentiment polarity lexicon of 1,195 nagamese words and use these to build features along with additional features for supervised machine learning techniques using Na"ive Bayes and Support Vector Machines. Keywords: Nagamese, NLP, sentiment analysis, machine learning
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。