arXiv:2605.08600cs.CL2026-05

构建10万+哈萨克斯坦电影评论数据集,支持多语言情感分析。

100,000+ Movie Reviews from Kazakhstan: Russian, Kazakh, and Code-Switched Texts

  • 采集2001-2025年10万+电影评论,含俄语、哈萨克语及混用文本。
  • 基于11,309条带评分评论,验证了预训练模型在情感分类中的优势。
  • 适合多语言自然语言处理、中亚语料研究者使用。

我们发布了一个公开的语料库,包含从kino.kz收集的100,502条哈萨克斯坦电影评论,时间跨度为2001-2025年,涵盖4,943部独特影片。该数据集为多语言,主要由俄语评论构成,同时包含哈萨克语和代码切换文本。所有评论均经人工标注语言与情感极性,其中11,309条评论附有用户明确评分。我们定义了两个情感任务:三分类极性识别和五分类评分识别,并将经典词袋/TF-IDF基线与多语言Transformer模型(mBERT、XLM-RoBERTa、RemBERT)进行对比。实验结果表明,变压器模型在极性分类任务上始终优于传统基线;而在受泄漏控制的评分分类任务中,由于类别严重失衡及相邻评分等级间差异细微,仍具挑战性。

原文摘要 · Abstract (English)

We present a new publicly available corpus of 100,502 movie reviews from Kazakhstan collected from kino.kz, spanning 2001-2025 and covering 4,943 unique titles. The dataset is multilingual, consisting mainly of Russian reviews alongside Kazakh and code-switched texts. Reviews are manually annotated for language and sentiment polarity, and 11,309 reviews additionally contain explicit user-provided ratings. We define two sentiment tasks -- three-way polarity classification and five-class score classification -- and benchmark classical BoW/TF-IDF baselines against multilingual transformer models (mBERT, XLM-RoBERTa, RemBERT). Experimental results show that transformer models consistently outperform classical baselines on polarity classification, while score classification remains challenging under leakage-controlled evaluation due to severe class imbalance and subtle distinctions between adjacent rating levels.

多语言情感分析语料库哈萨克斯坦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。