arXiv:2606.14068cs.CL2026-06被引 1

测试大模型对男女行为的道德评判差异,发现对男性更严厉。

Harsher on Male? Evaluating LLMs on Gender-Asymmetric Moral Framing Across Diverse Conflict Scenarios

论文配图:Harsher on Male? Evaluating LLMs on Gender-Asymmetric Moral Framing Across Diverse Conflict Scenarios
图 1 · 摘自论文原文
  • 构建1298个性别镜像场景,对比同罪不同性别的模型反应
  • 男性犯错被更多惩罚和归责,女性则更获同情与疗愈式回应
  • 现象跨模型、规模、推理方式均存在,适用于伦理评估研究

现有关于大模型性别偏见的研究多集中于刻板印象、职业关联或显性有害输出。本文探究大模型在男性与女性作为行为人时,是否对相同负面行为采取一致的回应标准。我们提出GAMA-Bench,一个包含1,298个场景的性别镜像基准,覆盖亲密关系与公共社会冲突。通过受控网格与跨模型评审构建性别中立的不当行为模板,并生成配对的第一人称提示,包含匹配的性别与角色引用变化。设计结构化响应框架,测量模型在惩罚、共情、升级、指导与归责方面的分配。在10个代表性大模型上的实验表明,存在一致的男性不利不对称:男性行为人获得更严厉、升级化与归责导向的表述,而女性则更多呈现治疗性与共情式回应。该模式在不同模型家族、场景类别、模型规模及显式思维推理中均持续存在。

原文摘要 · Abstract (English)

Existing studies on gender bias in LLMs have largely focused on stereotypes, occupational associations, or explicit harmful outputs. In this work, we ask whether LLMs apply consistent response standards to the same negative behavior under matched male-actor and female-actor conditions. We introduce GAMA-Bench, a gender-mirrored benchmark of 1,298 scenarios covering intimate relationship and public social conflicts. It constructs gender-neutral misconduct templates through controlled grids and cross-model review, then compiles them into paired first-person prompts with matched actor-gender and role-reference variations. We further design a structured response-framing protocol to measure how models allocate punishment, empathy, escalation, instruction, and blame. Experiments on 10 representative LLMs reveal a consistent male-disadvantaging asymmetry: male actors receive more punitive, escalatory, and blame-centered framing, whereas female actors receive more therapeutic and empathy-oriented framing for the same misconduct. Further analyses show that this pattern persists across model families, scenario tracks, model scale, and explicit thinking-style reasoning. The official code is available at https://github.com/xufeiqiong/GAMA-Bench.

大模型伦理性别偏见道德判断基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。