构建韩语细粒度情感分析数据集,提升电商评论情感识别精度
SSP-based construction of evaluation-annotated data for fine-grained aspect-based sentiment analysis
- 用半自动符号传播法标注数据,构建细粒度情感分析语料库
- 在电商场景中识别观点-值对,准确率F1达0.88~0.90
- 适合做韩国语情感分析、电商评论挖掘的研究者使用
本文构建了韩语评价标注语料库(Evaluation Annotated Dataset, EVAD),用于扩展面向电商评论的细粒度方面情感分析(ABSA)。通过半自动符号传播(SSP)方法,基于有限状态转换器(FST)构建语言资源,实现对电商领域中情感与非情感语言模式的精准标注。该方法不仅包含主题和方面,还引入方面值,并按单值、二值、多值分类,以更细致地捕捉用户观点。在评估中,基于KoBERT和KcBERT模型在该数据集上训练,对方面值对的识别分别达到F1 0.88和F1 0.90,表现稳健。
原文摘要 · Abstract (English)
We report the construction of a Korean evaluation-annotated corpus, hereafter called 'Evaluation Annotated Dataset (EVAD)', and its use in Aspect-Based Sentiment Analysis (ABSA) extended in order to cover e-commerce reviews containing sentiment and non-sentiment linguistic patterns. The annotation process uses Semi-Automatic Symbolic Propagation (SSP). We built extensive linguistic resources formalized as a Finite-State Transducer (FST) to annotate corpora with detailed ABSA components in the fashion e-commerce domain. The ABSA approach is extended, in order to analyze user opinions more accurately and extract more detailed features of targets, by including aspect values in addition to topics and aspects, and by classifying aspectvalue pairs depending whether values are unary, binary, or multiple. For evaluation, the KoBERT and KcBERT models are trained on the annotated dataset, showing robust performances of F1 0.88 and F1 0.90, respectively, on recognition of aspect-value pairs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。