arXiv:2502.20046cs.CLcs.AI2025-02被引 4

首个波兰语方面-情感三元组抽取数据集,助力多语言情感分析

Polish-ASTE: Aspect-Sentiment Triplet Extraction Datasets for Polish

  • 构建波兰语酒店与产品评论的三元组数据集
  • 在两个新数据集上验证模型性能,证明任务挑战性
  • 适配波兰语大模型,支持跨语言研究

方面-情感三元组抽取(ASTE)是情感分析中最具挑战性的任务之一,旨在构建包含方面、情感极性及支撑该极性的观点短语的三元组。尽管该任务日益受到关注,且已有多种机器学习方法提出,但相关数据集仍然稀缺,尤其缺乏斯拉夫语系的数据集。本文首次发布两个面向波兰语的ASTE数据集,涵盖酒店和商品评论。我们使用两种ASTE方法结合两种波兰语大模型进行实验,评估模型表现并验证数据集难度。新数据集采用与英文数据集相同的文件格式,可在宽松许可下自由使用,便于未来研究开展。

原文摘要 · Abstract (English)

Aspect-Sentiment Triplet Extraction (ASTE) is one of the most challenging and complex tasks in sentiment analysis. It concerns the construction of triplets that contain an aspect, its associated sentiment polarity, and an opinion phrase that serves as a rationale for the assigned polarity. Despite the growing popularity of the task and the many machine learning methods being proposed to address it, the number of datasets for ASTE is very limited. In particular, no dataset is available for any of the Slavic languages. In this paper, we present two new datasets for ASTE containing customer opinions about hotels and purchased products expressed in Polish. We also perform experiments with two ASTE techniques combined with two large language models for Polish to investigate their performance and the difficulty of the assembled datasets. The new datasets are available under a permissive licence and have the same file format as the English datasets, facilitating their use in future research.

情感分析三元组抽取波兰语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。