arXiv:2412.05964cs.CL2024-12被引 1

梳理十年土耳其语情感分析数据集与工具表现差异。

A Cross-Validation Study of Turkish Sentiment Analysis Datasets and Tools

  • 系统整理31篇研究中的23个土耳其语数据集,构建分类地图。
  • 测试主流工具在不同数据集上表现,发现文本特征影响显著。
  • 适合从事多语言情感分析或土耳其语NLP研究者参考。

近年来,情感分析日益重要,促使研究人员探索多种语言的数据集,包括土耳其语。然而,土耳其语数据集有限,导致其在不同研究中被反复使用,结果各异。为此,我们对2012至2022年间发表的研究进行了严格审查,共列出31项研究,并从公开来源及邮件请求中收集了23个在这些研究中使用的土耳其语数据集。我们采用分类体系对这31项研究进行标注,绘制出过去十年土耳其语情感分析数据集的分布图谱。此外,我们在这些主流土耳其语数据集上运行了最先进的情感分析工具,并分析了其性能表现。结果表明,情感分析工具的表现显著依赖于目标文本的特征。本研究促进了对土耳其语情感分析更深入的理解。

原文摘要 · Abstract (English)

In recent years, sentiment analysis has gained increasing significance, prompting researchers to explore datasets in various languages, including Turkish. However, the limited availability of Turkish datasets has led to their multifaceted usage in different studies, yielding diverse outcomes. To overcome this challenge, a rigorous review was conducted of research articles published between 2012 and 2022. 31 studies were listed, and 23 Turkish datasets obtained from publicly available sources and email requests used in these studies were collected. We labeled these 31 studies using a taxonomy. We provide a map of sentiment analysis datasets according to this taxonomy in Turkish over 10 years. Moreover, we run state-of-the-art sentiment analysis tools on these datasets and analyzed performance across popular Turkish sentiment datasets. We observed that the performance of the sentiment analysis tools significantly depends on the characteristics of the target text. Our study fosters a more nuanced understanding of sentiment analysis in the Turkish language.

情感分析土耳其语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。