arXiv:2509.05719cs.CL2025-09综述

分析波斯语主观任务研究现状,揭示数据稀缺与质量不足问题

Exploring Subjective Tasks in Farsi: A Survey Analysis and Evaluation of Language Models

  • 梳理110篇波斯语主观任务论文,发现公开数据集严重缺失
  • 现有数据缺乏年龄、性别等关键人口统计信息,影响模型准确性
  • 模型在少数可用数据集上表现极不稳定,凸显数据量不足

尽管波斯语拥有超过1.27亿使用者和超过130万篇维基百科文章,被视为中等资源语言,但深入分析发现其在主观任务(情感分析、情绪分析、毒性检测)方面仍面临严峻挑战。我们系统回顾了110篇相关论文,发现公开数据集极度匮乏。现有数据集普遍缺少年龄、性别等关键人口统计特征,难以准确建模语言中的主观性。在仅有的几个数据集上评估模型时,结果在不同数据集和模型间波动剧烈,表明当前数据规模不足以支撑有效提升波斯语NLP性能。

原文摘要 · Abstract (English)

Given Farsi's speaker base of over 127 million people and the growing availability of digital text, including more than 1.3 million articles on Wikipedia, it is considered a middle-resource language. However, this label quickly crumbles when the situation is examined more closely. We focus on three subjective tasks (Sentiment Analysis, Emotion Analysis, and Toxicity Detection) and find significant challenges in data availability and quality, despite the overall increase in data availability. We review 110 publications on subjective tasks in Farsi and observe a lack of publicly available datasets. Furthermore, existing datasets often lack essential demographic factors, such as age and gender, that are crucial for accurately modeling subjectivity in language. When evaluating prediction models using the few available datasets, the results are highly unstable across both datasets and models. Our findings indicate that the volume of data is insufficient to significantly improve a language's prospects in NLP.

波斯语主观任务数据稀缺

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。