arXiv:2512.23732cs.CLcs.AI2025-12ACL

用专家辩论机制提升性别歧视检测的准确率和鲁棒性

When in Doubt, Consult: Expert Debate for Sexism Detection via Confidence-Based Routing

  • 通过动态路由将简单样本与复杂样本分离,仅对模糊案例启用多专家推理
  • 在多个公开数据集上实现最高F1提升4.48%,显著改善小样本与噪声标签问题
  • 适合需要高精度、可解释性判断的敏感内容审核场景

在线性别歧视日益以微妙、依赖语境的形式出现,传统方法难以捕捉。其判断涉及语言、心理、法律和文化多重维度,导致标注数据存在矛盾信号。结合标签稀缺与类别不平衡,使模型决策边界不稳定,难以识别较隐蔽的伤害形式。为此,我们提出两阶段框架:第一阶段采用类平衡焦点损失、类别感知批量处理及事后阈值校准,首次应用于该领域以缓解标签不均衡与噪声;第二阶段引入动态路由机制,区分明确与模糊案例,并通过新型协同专家判断(CEJ)模块,让多个角色身份进行推理并由裁判模型整合结论。实验表明,本方法在多个公开基准上优于现有模型,在EDOS任务A和B上分别获得+4.48%和+1.30%的F1提升,以及在EXIST 2025 Task 1.1中实现+2.79%的改进。

原文摘要 · Abstract (English)

Online sexism increasingly appears in subtle, context-dependent forms that evade traditional detection methods. Its interpretation often depends on overlapping linguistic, psychological, legal, and cultural dimensions, which produce mixed and sometimes contradictory signals in annotated datasets. These inconsistencies, combined with label scarcity and class imbalance, result in unstable decision boundaries and cause fine-tuned models to overlook subtler, underrepresented forms of harm. To address these challenges, we propose a two-stage framework that unifies (i) targeted training procedures to better regularize supervision to scarce and noisy data with (ii) selective, reasoning-based inference to handle ambiguous or borderline cases. First, we stabilize the training combining class-balanced focal loss, class-aware batching, and post-hoc threshold calibration, strategies for the firs time adapted for this domain to mitigate label imbalance and noisy supervision. Second, we bridge the gap between efficiency and reasoning with a a dynamic routing mechanism that distinguishes between unambiguous instances and complex cases requiring a deliberative process. This reasoning process results in the novel Collaborative Expert Judgment (CEJ) module which prompts multiple personas and consolidates their reasoning through a judge model. Our approach outperforms existing approaches across several public benchmarks, with F1 gains of +4.48% and +1.30% on EDOS Tasks A and B, respectively, and a +2.79% improvement in ICM on EXIST 2025 Task 1.1.

性别歧视检测多专家推理动态路由弱监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。