arXiv:2501.06265cs.CL2025-01被引 2

构建了171篇希腊大选演讲的多标注数据集,支持政治分析与模型训练

AgoraSpeech: A multi-annotated comprehensive dataset of political discourse through the lens of humans and AI

  • 用AI初标+人工复核方式,对每段演讲标注6类NLP任务
  • 包含6个政党共171篇2023年希腊选举演讲,覆盖语义、情感、极化等维度
  • 适合政治学者、记者及模型开发者用于研究或训练大模型

政治话语数据集对于获取政治洞察、分析传播策略或社会现象至关重要。尽管已有大量政治话语语料库,但高质量、多维度标注的数据集仍十分稀缺,主要受限于人工成本高、跨学科要求严以及对修辞策略和意识形态背景的细致理解需求。本文提出AgoraSpeech,一个精心构建的高质量数据集,涵盖2023年希腊全国选举期间六个政党的171篇政治演讲。数据集对每段文本进行了六项自然语言处理(NLP)任务的标注:文本分类、主题识别、情感分析、命名实体识别、极化检测和民粹主义识别。采用两阶段标注流程:先由ChatGPT生成初步标注,再通过严格的人工闭环验证。该数据集最初用于选举前的案例研究,以提供政治洞见,但具有广泛适用性,可作为政治与社会科学家、记者或数据科学家的重要信息来源,并可用于基准测试和大语言模型(LLMs)的微调。

原文摘要 · Abstract (English)

Political discourse datasets are important for gaining political insights, analyzing communication strategies or social science phenomena. Although numerous political discourse corpora exist, comprehensive, high-quality, annotated datasets are scarce. This is largely due to the substantial manual effort, multidisciplinarity, and expertise required for the nuanced annotation of rhetorical strategies and ideological contexts. In this paper, we present AgoraSpeech, a meticulously curated, high-quality dataset of 171 political speeches from six parties during the Greek national elections in 2023. The dataset includes annotations (per paragraph) for six natural language processing (NLP) tasks: text classification, topic identification, sentiment analysis, named entity recognition, polarization and populism detection. A two-step annotation was employed, starting with ChatGPT-generated annotations and followed by exhaustive human-in-the-loop validation. The dataset was initially used in a case study to provide insights during the pre-election period. However, it has general applicability by serving as a rich source of information for political and social scientists, journalists, or data scientists, while it can be used for benchmarking and fine-tuning NLP and large language models (LLMs).

政治话语多任务标注大语言模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。