arXiv:2509.07889cs.CL2025-09

用高效微调与投票机制,提升中文文本性别偏见检测与缓解效果。

From Detection to Mitigation: Addressing Gender Bias in Chinese Texts via Efficient Tuning and Voting-Based Rebalancing

  • 基于LoRA的高效微调,快速适配大模型进行偏见检测。
  • 多专家投票+多温度采样,提升检测准确率与偏见缓解能力。
  • 适合关注中文NLP公平性与可控生成的研究者使用。

本文提出团队在NLPCC-2025共享任务7中的解决方案,聚焦中文句子级性别偏见的检测与缓解。该任务旨在通过自动检测、分类和缓解性别偏见,提升自然语言生成的公平性与可控性。我们采用基于大语言模型的微调方法,利用低秩适配(LoRA)实现高效参数更新;在数据处理方面,构建更均衡的训练集以缓解类别不平衡问题,并引入多源异构样本以增强模型泛化能力。对于检测与分类子任务,采用多专家模型的多数投票策略提升性能;为改善偏见生成的检测与缓解效果,设计多温度采样机制以捕捉偏见表达风格的多样性。实验结果表明,该方法在偏见检测、分类与缓解任务中均表现有效,最终平均得分47.90%,在共享任务中排名第四。

原文摘要 · Abstract (English)

This paper presents our team's solution to Shared Task 7 of NLPCC-2025, which focuses on sentence-level gender bias detection and mitigation in Chinese. The task aims to promote fairness and controllability in natural language generation by automatically detecting, classifying, and mitigating gender bias. To address this challenge, we adopt a fine-tuning approach based on large language models (LLMs), efficiently adapt to the bias detection task via Low-Rank Adaptation (LoRA). In terms of data processing, we construct a more balanced training set to alleviate class imbalance and introduce heterogeneous samples from multiple sources to enhance model generalization. For the detection and classification sub-tasks, we employ a majority voting strategy that integrates outputs from multiple expert models to boost performance. Additionally, to improve bias generation detection and mitigation, we design a multi-temperature sampling mechanism to capture potential variations in bias expression styles. Experimental results demonstrate the effectiveness of our approach in bias detection, classification, and mitigation. Our method ultimately achieves an average score of 47.90%, ranking fourth in the shared task.

性别偏见中文NLP大模型微调公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。