构建德国各级议会演讲语料库,含丰富元数据支持政治话语分析
SpeakGer: A meta-data enriched speech corpus of German state and federal parliaments
- 整合1947-2023年德国16州及联邦议院1080万条演讲,含发言者政党、年龄等元数据
- 揭示不同政党随时间的主题分布差异,发现疫情期间各党言论情绪显著分化
- 适合政治科学、计算社会科学领域研究者,尤其关注话语与身份关联的分析
自然语言处理在政治文本和演讲分析中的应用日益重要,但多数语料缺乏发言者所属政党、年龄、选区等关键元信息,难以支撑细粒度研究。为此,我们构建了SpeakGer语料库,涵盖1947至2023年德国16个联邦州及联邦议院的议会辩论内容,共包含10,806,105条演讲记录。该数据集不仅包含发言者所属政党、年龄、选区及其政党政治倾向等元数据,还收录了观众反应信息,支持更深入的定量分析。我们进一步提供了三项探索性分析:不同政党随时间的主题分布变化、平均发言者年龄演变趋势,以及针对新冠疫情下各党言论的情感分析结果。
原文摘要 · Abstract (English)
The application of natural language processing on political texts as well as speeches has become increasingly relevant in political sciences due to the ability to analyze large text corpora which cannot be read by a single person. But such text corpora often lack critical meta information, detailing for instance the party, age or constituency of the speaker, that can be used to provide an analysis tailored to more fine-grained research questions. To enable researchers to answer such questions with quantitative approaches such as natural language processing, we provide the SpeakGer data set, consisting of German parliament debates from all 16 federal states of Germany as well as the German Bundestag from 1947-2023, split into a total of 10,806,105 speeches. This data set includes rich meta data in form of information on both reactions from the audience towards the speech as well as information about the speaker's party, their age, their constituency and their party's political alignment, which enables a deeper analysis. We further provide three exploratory analyses, detailing topic shares of different parties throughout time, a descriptive analysis of the development of the age of an average speaker as well as a sentiment analysis of speeches of different parties with regards to the COVID-19 pandemic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。