首个面向乌兹别克语和维吾尔语的细粒度情感四元组数据集
LASQ: A Low-resource Aspect-based Sentiment Quadruple Extraction Dataset

- 构建乌兹别克语与维吾尔语的细粒度情感四元组数据集
- 在低资源语言上实现优于基线模型的抽取效果
- 适合低资源语言情感分析研究者使用
近年来,基于方面的情感分析(ABSA)发展迅速并展现出强大的实际应用价值。然而,现有研究与基准测试主要集中于高资源语言,导致低资源语言中的细粒度情感抽取仍处于探索阶段。为填补这一空白,我们构建了首个面向低资源语言的细粒度情感四元组数据集LASQ,涵盖乌兹别克语和维吾尔语两种语言。该数据集包含目标-方面-观点-情感四元组抽取任务。为促进后续研究,我们设计了一种融合句法知识的网格标注模型,通过自研的句法知识嵌入模块(SKEM)将词性(POS)和依存关系知识融入模型,有效缓解黏着语系导致的词汇稀疏问题。在LASQ上的实验表明,该方法持续优于多个竞争性基线模型,验证了数据集的有效性及建模方法的优越性。
原文摘要 · Abstract (English)
In recent years, aspect-based sentiment analysis (ABSA) has made rapid progress and shown strong practical value. However, existing research and benchmarks are largely concentrated on high-resource languages, leaving fine-grained sentiment extraction in low-resource languages under-explored. To address this gap, we constructed the first Low-resource languages Aspect-based Sentiment Quadruple dataset, named LASQ, which includes two low-resource languages: Uzbek and Uyghur. Secondly, it includes a fine-grained target-aspect-opinion-sentiment quadruple extraction task. To facilitate future research, we designed a grid-tagging model that integrates syntactic knowledge. This model incorporates part-of-speech (POS) and dependency knowledge into the model through our designed Syntax Knowledge Embedding Module (SKEM), thereby alleviating the lexical sparsity problem caused by agglutinative languages. Experiments on LASQ demonstrate consistent gains over competitive baselines, validating both the dataset's utility and the effectiveness of the proposed modeling approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。