arXiv:2510.08214cs.CL2025-10

构建多语言细粒度情感数据集,用于分析新冠疫情期间公众情绪变化。

SenWave: A Fine-Grained Multi-Language Sentiment Analysis Dataset Sourced from COVID-19 Tweets

  • 基于英文、阿拉伯文等五种语言的10万条标注推文构建数据集
  • 涵盖十类情感标签,包含超1亿条未标注新冠推文用于研究
  • 适合作为复杂事件下细粒度情感分析的基准数据集

新冠疫情全球蔓延凸显了理解公众情绪反应的重要性。尽管已有大量新冠相关数据集(部分达百亿量级),但标注数据稀缺且情感标签粗粒度或不准确的问题仍存。本文提出SenWave,一个面向新冠推文的细粒度多语言情感分析数据集,覆盖英语、阿拉伯语等五种语言,包含每种语言1万条标注推文,以及3万条由英文推文翻译的西语、法语和意大利语推文。此外,还收录超过1.05亿条未标注推文,采集自不同疫情阶段。为实现精准细粒度分类,我们使用标注数据微调预训练的Transformer模型。研究深入分析了多语言、多国别、多话题下的情绪演变,揭示了时间维度上的显著变化。同时评估了该数据集与ChatGPT的兼容性,验证其在多种应用中的鲁棒性与通用性。数据集及代码已开源。本工作有望推动自然语言处理领域对复杂事件的细粒度情感研究,促进更深入的理解与创新。

原文摘要 · Abstract (English)

The global impact of the COVID-19 pandemic has highlighted the need for a comprehensive understanding of public sentiment and reactions. Despite the availability of numerous public datasets on COVID-19, some reaching volumes of up to 100 billion data points, challenges persist regarding the availability of labeled data and the presence of coarse-grained or inappropriate sentiment labels. In this paper, we introduce SenWave, a novel fine-grained multi-language sentiment analysis dataset specifically designed for analyzing COVID-19 tweets, featuring ten sentiment categories across five languages. The dataset comprises 10,000 annotated tweets each in English and Arabic, along with 30,000 translated tweets in Spanish, French, and Italian, derived from English tweets. Additionally, it includes over 105 million unlabeled tweets collected during various COVID-19 waves. To enable accurate fine-grained sentiment classification, we fine-tuned pre-trained transformer-based language models using the labeled tweets. Our study provides an in-depth analysis of the evolving emotional landscape across languages, countries, and topics, revealing significant insights over time. Furthermore, we assess the compatibility of our dataset with ChatGPT, demonstrating its robustness and versatility in various applications. Our dataset and accompanying code are publicly accessible on the repository\footnote{https://github.com/gitdevqiang/SenWave}. We anticipate that this work will foster further exploration into fine-grained sentiment analysis for complex events within the NLP community, promoting more nuanced understanding and research innovations.

情感分析多语言新冠数据细粒度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。