arXiv:2607.05259cs.CLcs.LG2026-07

构建首个僧伽罗语细粒度情感分析数据集,助力低资源语言NLP研究

SalAngaBhava: A Sinhala Market Dataset for Aspect-based Sentiment Analysis

论文配图:SalAngaBhava: A Sinhala Market Dataset for Aspect-based Sentiment Analysis
图 1 · 摘自论文原文
  • 手工标注僧伽罗语产品评论中的观点词与情感极性
  • 涵盖多领域观点项,数据结构清晰且平衡性良好
  • 为斯里兰卡本地语言和低资源场景提供基准数据

情感分析是自然语言处理的核心领域,广泛应用于现实与科研场景。在高资源语言中,研究已从句子级情感转向更精细的方面级情感分析(ABSA),但需依赖带有观点项与对应情感标注的数据集。目前,低资源语言如僧伽罗语(Sinhala)仍缺乏此类数据。本文提出 SalAngaBhava,首个僧伽罗语方面级情感分析数据集,包含来自用户生成评论和评论区的手动标注数据,涵盖多个领域的真实观点项及其情感标签(正面、负面、中性)。数据按严格标注指南构建,确保一致性和质量。分析表明该数据集结构良好、类别分布均衡,可作为僧伽罗语自然语言处理及低资源情感分析研究的基准。

原文摘要 · Abstract (English)

Sentiment analysis has been a primary domain under Natural Language Processing (NLP) from its inception as it plays a vital role in both real-world and research applications. In high-resource languages, this has been extended a step further, and instead of predicting sentiment at the sentence level, models have been developed to detect more fine-grained sentiments at aspect level. However, in order to conduct this fine-grained Aspect-based Sentiment Analysis (ABSA), datasets annotated with aspects and sentiments toward the said aspects is required. Such datasets are lacking for low-resources languages among which, we can count Sinhala, an Indo-Aryan languages used primarily in Sri Lanka. In this work, we introduce, SalAngaBhava, a new Sinhala Aspect-based Sentiment Analysis dataset which contains Sinhala product reviews that are manually labeled with aspect terms and the associated sentiments (positive, negative, neutral). The data was collected from domain-relevant sources such as user-generated reviews and comments, and was annotated following carefully defined guidelines to ensure consistency and quality. The dataset consists of sentences and aspect-sentiment pairs, encompassing a considerable range of aspects from several domains. The analysis confirms that the dataset is well-structured and sufficiently balanced for ABSA research. This dataset can be used as a benchmark and facilitates further studies related to Sinhala natural language processing, and low-resource sentiment analysis tasks.

情感分析低资源语言僧伽罗语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。