arXiv:2508.13169cs.CLcs.CY2025-08

通过角色级分析,系统性减少新闻语料中的性别偏见

Fair Play in the Newsroom: Actor-Based Filtering Gender Discrimination in Text Corpora

  • 基于对话上下文,从情感、句法主动性和引语风格多维度检测性别偏见
  • 在德语报纸语料中实现性别平衡,保留原始内容核心特征
  • 适合关注语言公平性与数据伦理的研究者使用

语言语料是大多数自然语言处理研究的基础,但常复制结构性不平等。其中,角色代表性上的性别歧视会扭曲分析结果并加剧歧视性后果。本文提出一种以用户为中心、基于角色的流水线方法,用于检测和缓解大规模文本语料中的性别歧视。结合话语感知分析与情感、句法主动性、引语风格等指标,该方法支持细粒度审计与排除式平衡。应用于1980-2024年德语报纸语料taz2024full,结果生成了更性别均衡的数据集,同时保留了原始材料的核心动态。研究发现,结构性不对称可通过系统性过滤降低,但情感与表述层面的细微偏见仍存在。我们开源工具与报告,以支持基于话语的公平性审计与公正语料构建。

原文摘要 · Abstract (English)

Language corpora are the foundation of most natural language processing research, yet they often reproduce structural inequalities. One such inequality is gender discrimination in how actors are represented, which can distort analyses and perpetuate discriminatory outcomes. This paper introduces a user-centric, actor-level pipeline for detecting and mitigating gender discrimination in large-scale text corpora. By combining discourse-aware analysis with metrics for sentiment, syntactic agency, and quotation styles, our method enables both fine-grained auditing and exclusion-based balancing. Applied to the taz2024full corpus of German newspaper articles (1980-2024), the pipeline yields a more gender-balanced dataset while preserving core dynamics of the source material. Our findings show that structural asymmetries can be reduced through systematic filtering, though subtler biases in sentiment and framing remain. We release the tools and reports to support further research in discourse-based fairness auditing and equitable corpus construction.

性别偏见语料公平话语分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。