arXiv:2506.12574cs.CL2025-06综述被引 2

针对中文文本性别偏见,构建了带标注的语料库并发起自动消减挑战。

Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge

  • 构建32.9k句中文语料库,含人工修正的5.2k条偏见消除句。
  • 提出三项共享任务:检测、分类与自动化消除文本性别偏见。
  • 为资源稀缺语言如中文提供可落地的偏见评估与缓解工具。

随着自然语言处理中性别偏见问题日益受到关注,基于数据驱动的预训练语言模型常受有偏语料影响,这一现象在缺乏公平性计算语言资源的语言中尤为明显,如中文。为此,我们提出了一个面向中文性别偏见探测与缓解的语料库(CORGI-PM),包含32.9千条高质量标注句子,其标注方案专为中文语境设计。该语料库包含5.2千条具有性别偏见的句子及其由人工标注者重写的无偏版本。我们设立了三项共享任务,旨在推动文本性别偏见的自动化检测、分类与缓解。本文报告了2025年NLPCC会议上参与团队的相关结果与分析。

原文摘要 · Abstract (English)

As natural language processing for gender bias becomes a significant interdisciplinary topic, the prevalent data-driven techniques, such as pre-trained language models, suffer from biased corpus. This case becomes more obvious regarding those languages with less fairness-related computational linguistic resources, such as Chinese. To this end, we propose a Chinese cOrpus foR Gender bIas Probing and Mitigation (CORGI-PM), which contains 32.9k sentences with high-quality labels derived by following an annotation scheme specifically developed for gender bias in the Chinese context. It is worth noting that CORGI-PM contains 5.2k gender-biased sentences along with the corresponding bias-eliminated versions rewritten by human annotators. We pose three challenges as a shared task to automate the mitigation of textual gender bias, which requires the models to detect, classify, and mitigate textual gender bias. In the literature, we present the results and analysis for the teams participating this shared task in NLPCC 2025.

性别偏见中文NLP语料库自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。