融合多种嵌入与依存结构,提升细粒度情感分析中的方面词抽取效果。
An Aspect Extraction Framework using Different Embedding Types, Learning Models, and Dependency Structure
- 使用词与词性嵌入结合多种学习模型进行方面词提取。
- 引入依存树位置编码,显著提升模型对方面词位置的捕捉能力。
- 构建首个土耳其语方面抽取数据集,适合多语言情感分析研究者。
基于方面的情感分析近年来受到广泛关注,因其能为实体特定特征的情感表达提供细粒度洞察。其关键环节是方面词抽取,即从文本中识别并提取方面术语。有效的方面词抽取是实现精确方面级情感分析的基础。本文提出一种方面抽取框架,采用不同类型的词嵌入和词性标签嵌入,并结合多种学习模型。同时,提出基于依存句法分析输出的树形位置编码,以更好捕捉句子中方面词的位置。此外,在受控环境下通过机器翻译构建了一个新的土耳其语方面抽取数据集。在两个土耳其语数据集上的实验表明,所提模型多数情况下优于使用相同数据集的已有研究,且引入树形位置编码可进一步提升性能。
原文摘要 · Abstract (English)
Aspect-based sentiment analysis has gained significant attention in recent years due to its ability to provide fine-grained insights for sentiment expressions related to specific features of entities. An important component of aspect-based sentiment analysis is aspect extraction, which involves identifying and extracting aspect terms from text. Effective aspect extraction serves as the foundation for accurate sentiment analysis at the aspect level. In this paper, we propose aspect extraction models that use different types of embeddings for words and part-of-speech tags and that combine several learning models. We also propose tree positional encoding that is based on dependency parsing output to capture better the aspect positions in sentences. In addition, a new aspect extraction dataset is built for Turkish by machine translating an English dataset in a controlled setting. The experiments conducted on two Turkish datasets showed that the proposed models mostly outperform the studies that use the same datasets, and incorporating tree positional encoding increases the performance of the models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。