构建欧洲议会议程数据集,用大模型高效标注政策主题
Supercharging Agenda Setting Research: The ParlaCAP Dataset of 28 European Parliaments and a Scalable Multilingual LLM-Based Classification
- 用大模型自动标注800万条议会发言,生成多语言分类器
- 模型准确率媲美人工标注,超越旧有基于人工数据的分类器
- 适合研究政治关注、性别差异与跨国家政策比较的学者
本文提出ParlaCAP,一个涵盖28个欧洲议会的大型多语言议程分析数据集,基于超过800万条演讲的ParlaMint语料库。采用教师-学生框架:高性能大语言模型(LLM)对领域内数据进行标注,再以这些标注数据微调多语言编码模型,实现可扩展的自动化标注。结果表明,该方法生成的分类器精准适配目标领域,且大模型与人工标注者的一致性达到人类标注者间水平,优于以往在非领域人工标注数据上训练的CAP分类器。除标准的CAP标注外,数据集还包含丰富的发言人和政党元数据,以及由ParlaSent多语言模型生成的情感预测,支持跨国政治关注度与代表性对比研究。通过三个应用案例展示了其分析潜力:政策议题关注分布、演讲情感模式及性别差异下的政策关注度。
原文摘要 · Abstract (English)
This paper introduces ParlaCAP, a large-scale dataset for analyzing parliamentary agenda setting across Europe, and proposes a cost-effective method for building domain-specific policy topic classifiers. Applying the Comparative Agendas Project (CAP) schema to the multilingual ParlaMint corpus of over 8 million speeches from 28 parliaments of European countries and autonomous regions, we follow a teacher-student framework in which a high-performing large language model (LLM) annotates in-domain training data and a multilingual encoder model is fine-tuned on these annotations for scalable data annotation. We show that this approach produces a classifier tailored to the target domain. Agreement between the LLM and human annotators is comparable to inter-annotator agreement among humans, and the resulting model outperforms existing CAP classifiers trained on manually-annotated but out-of-domain data. In addition to the CAP annotations, the ParlaCAP dataset offers rich speaker and party metadata, as well as sentiment predictions coming from the ParlaSent multilingual transformer model, enabling comparative research on political attention and representation across countries. We illustrate the analytical potential of the dataset with three use cases, examining the distribution of parliamentary attention across policy topics, sentiment patterns in parliamentary speech, and gender differences in policy attention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。