融合PhoBERT-V2与SentiWordNet提升越南语情感分析效果
Expanding Vietnamese SentiWordNet to Improve Performance of Vietnamese Sentiment Analysis Models
- 结合预训练模型PhoBERT-V2与词典SentiWordNet进行情感分类
- 在VLSP 2016和AIVIVN 2019数据集上表现优于现有模型
- 适合需要高精度越南语情感分析的研究与应用
情感分析是自然语言处理中至关重要的任务,旨在训练机器学习模型以根据观点极性对文本进行分类。通过微调,预训练语言模型(PLMs)可直接应用于下游任务,无需从头训练。近年来,预训练的PhoBERT-V2模型已成为越南语情感分析的最先进方法,其基于RoBERTa架构,优化了BERT的预训练方式,提升了鲁棒性。本文提出一种新方法,将PhoBERT-V2与SentiWordNet相结合,用于越南语评论的情感分析。该模型利用PhoBERT-V2对越南语进行深度语义理解,并借助专为情感分类设计的词典SentiWordNet增强极性识别能力。在VLSP 2016与AIVIVN 2019数据集上的实验结果表明,所提系统在性能上显著优于其他对比模型。
原文摘要 · Abstract (English)
Sentiment analysis is one of the most crucial tasks in Natural Language Processing (NLP), involving the training of machine learning models to classify text based on the polarity of opinions. Pre-trained Language Models (PLMs) can be applied to downstream tasks through fine-tuning, eliminating the need to train the model from scratch. Specifically, PLMs have been employed for Sentiment Analysis, a process that involves detecting, analyzing, and extracting the polarity of text sentiments. Numerous models have been proposed to address this task, with pre-trained PhoBERT-V2 models standing out as the state-of-the-art language models for Vietnamese. The PhoBERT-V2 pre-training approach is based on RoBERTa, optimizing the BERT pre-training method for more robust performance. In this paper, we introduce a novel approach that combines PhoBERT-V2 and SentiWordnet for Sentiment Analysis of Vietnamese reviews. Our proposed model utilizes PhoBERT-V2 for Vietnamese, offering a robust optimization for the prominent BERT model in the context of Vietnamese language, and leverages SentiWordNet, a lexical resource explicitly designed to support sentiment classification applications. Experimental results on the VLSP 2016 and AIVIVN 2019 datasets demonstrate that our sentiment analysis system has achieved excellent performance in comparison to other models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。