针对印度社交媒体中的印英混用文本,提出高效情感分析框架
Code-Mix Sentiment Analysis on Hinglish Tweets
- 用mBERT结合子词分词处理印英混用语料
- 在Hinglish数据集上实现91.3%准确率
- 适合需要本地化品牌舆情监控的团队
印度品牌舆情监测正面临日益增长的挑战,源于推特等平台用户生成内容中广泛使用的印英混用语言(Hinglish)。传统基于单语数据的自然语言处理模型难以解析此类混合语言的句法与语义复杂性,导致情感分析不准确,市场洞察失真。为此,本文提出一种专为Hinglish推文设计的高性能情感分类框架。方法上,通过微调mBERT(多语言BERT),利用其多语言能力更好地理解印度社交媒体的语言多样性。关键创新在于采用子词分词机制,有效应对罗马化拼写差异、俚语及未登录词等问题。本研究提供了一个可投入生产的AI解决方案,并在低资源、代码混用环境下建立了坚实基准。
原文摘要 · Abstract (English)
The effectiveness of brand monitoring in India is increasingly challenged by the rise of Hinglish--a hybrid of Hindi and English--used widely in user-generated content on platforms like Twitter. Traditional Natural Language Processing (NLP) models, built for monolingual data, often fail to interpret the syntactic and semantic complexity of this code-mixed language, resulting in inaccurate sentiment analysis and misleading market insights. To address this gap, we propose a high-performance sentiment classification framework specifically designed for Hinglish tweets. Our approach fine-tunes mBERT (Multilingual BERT), leveraging its multilingual capabilities to better understand the linguistic diversity of Indian social media. A key component of our methodology is the use of subword tokenization, which enables the model to effectively manage spelling variations, slang, and out-of-vocabulary terms common in Romanized Hinglish. This research delivers a production-ready AI solution for brand sentiment tracking and establishes a strong benchmark for multilingual NLP in low-resource, code-mixed environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。