arXiv:2508.08125cs.CL2025-08被引 8

首个面向复杂情感分析的捷克语数据集,支持多任务统一标注。

Czech Dataset for Complex Aspect-Based Sentiment Analysis Tasks

  • 构建统一标注格式,支持目标-方面类别等复杂任务
  • 3100条餐厅评论,标注一致率高达90%
  • 适配跨语言对比,助力多语言情感分析研究

本文提出首个面向复杂方面情感分析(ABSA)的捷克语数据集,包含3100条人工标注的餐厅评论。该数据集基于早期仅支持基础任务(如方面词抽取或极性检测)的捷克语数据集,但新增了目标-方面类别检测等高级任务。采用与SemEval-2016一致的标注格式,便于跨语言比较和评估。标注由两名训练有素的标注员完成,互评一致性达约90%。此外,我们还提供了2400万条无标注评论,可用于无监督学习。本文展示了多种Transformer模型在单语场景下的基线性能,并进行了深入的错误分析。代码与数据集可免费用于非商业研究。

原文摘要 · Abstract (English)

In this paper, we introduce a novel Czech dataset for aspect-based sentiment analysis (ABSA), which consists of 3.1K manually annotated reviews from the restaurant domain. The dataset is built upon the older Czech dataset, which contained only separate labels for the basic ABSA tasks such as aspect term extraction or aspect polarity detection. Unlike its predecessor, our new dataset is specifically designed for more complex tasks, e.g. target-aspect-category detection. These advanced tasks require a unified annotation format, seamlessly linking sentiment elements (labels) together. Our dataset follows the format of the well-known SemEval-2016 datasets. This design choice allows effortless application and evaluation in cross-lingual scenarios, ultimately fostering cross-language comparisons with equivalent counterpart datasets in other languages. The annotation process engaged two trained annotators, yielding an impressive inter-annotator agreement rate of approximately 90%. Additionally, we provide 24M reviews without annotations suitable for unsupervised learning. We present robust monolingual baseline results achieved with various Transformer-based models and insightful error analysis to supplement our contributions. Our code and dataset are freely available for non-commercial research purposes.

情感分析捷克语数据集ABSA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。