arXiv:2506.05388cs.CL2025-06ACL被引 3

构建180万篇德语报纸文集,分析近40年性别偏见演变

taz2024full: Analysing German Newspapers for Gender Bias and Discrimination across Decades

  • 整合1980至2024年《时代周报》超180万篇文章,建立最大公开德语语料库
  • 发现男性在报道中持续被过度代表,近年呈现逐步平衡趋势
  • 提供可扩展分析框架,适合语言演化与媒体批判研究者使用

开放获取语料库对自然语言处理(NLP)和计算社会科学研究至关重要。然而,大规模德语资源仍有限,制约了对语言趋势与性别偏见等社会议题的研究。本文发布 taz2024full,迄今最大的公开德语报纸文章语料库,包含来自《时代周报》(taz)的180多万篇文本,覆盖1980至2024年。为展示其在偏见研究中的价值,我们分析了四十年间性别代表性变化:男性持续被过度代表,但近年来出现逐步趋近平衡的迹象。通过可扩展的结构化分析流程,本文为德语新闻文本中的角色提及、情感倾向与语言框架研究提供了基础。该语料库适用于历时语言分析、媒体批判研究等多种场景,且免费开放,旨在推动德语NLP领域更包容、可复现的研究。

原文摘要 · Abstract (English)

Open-access corpora are essential for advancing natural language processing (NLP) and computational social science (CSS). However, large-scale resources for German remain limited, restricting research on linguistic trends and societal issues such as gender bias. We present taz2024full, the largest publicly available corpus of German newspaper articles to date, comprising over 1.8 million texts from taz, spanning 1980 to 2024. As a demonstration of the corpus's utility for bias and discrimination research, we analyse gender representation across four decades of reporting. We find a consistent overrepresentation of men, but also a gradual shift toward more balanced coverage in recent years. Using a scalable, structured analysis pipeline, we provide a foundation for studying actor mentions, sentiment, and linguistic framing in German journalistic texts. The corpus supports a wide range of applications, from diachronic language analysis to critical media studies, and is freely available to foster inclusive and reproducible research in German-language NLP.

德语NLP性别偏见语料库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。