arXiv:2606.02214cs.CL2026-06

性别提示影响大模型决策,但模型常误判原因。

Do Gender Cues Affect LLM Value Trade-offs? Evidence from a Controlled Decision Benchmark

论文配图:Do Gender Cues Affect LLM Value Trade-offs? Evidence from a Controlled Decision Benchmark
图 1 · 摘自论文原文
  • 构建可控基准测试,仅改变角色性别配置。
  • 性别提示导致系统性决策翻转,尤其在高严重度情境下。
  • 模型常归因于无影响,实际却受性别暗示干扰。

大型语言模型在价值敏感决策场景中应用日益广泛,此时无关的人口统计学线索不应改变判断。本文构建了真实价值决策基准(RVDB),仅改变角色性别配置,而保持情景、价值对、角色、候选决策、价值距离与决策严重性不变。通过在七种模型上进行位置平衡评估,检验模型在性别扰动下是否保持决策不变性,以及其自我归因是否反映实际行为变化。结果发现,明确的性别提示会引发有界但系统的决策翻转,包括在要求模型报告性别是否影响选择的提示下仍出现翻转。跨性别角色互换显示女性提议决策存在一致性偏差,而模型多将翻转归因为无影响或其他非性别因素。进一步分析表明,性别效应集中在价值判断边界模糊处且在更严重决策情境下更显著,说明性别提示是局部边界偏移因素而非全局价值推理覆盖。价值排序总体稳定,但价值对权衡在不同性别角色配置下不均衡变化。结果表明,性别可实质性影响模型价值权衡,但自我归因却掩盖这一现象,呼吁超越解释性评估的可控行为审计。

原文摘要 · Abstract (English)

Large language models are increasingly used in value-sensitive decision settings, where irrelevant demographic cues should not alter judgments. We construct the Realistic Value Decision Benchmark (RVDB), a controlled benchmark that varies only the role-gender configuration while holding the scenario, ordered value pair, roles, candidate decisions, Value Distance, and Decision Severity fixed. Using a position-balanced evaluation across seven models, we test whether models preserve decision invariance under gender perturbations and whether their self-attributions reflect observed behavioral changes. We find that explicit gender cues induce bounded but systematic decision flips, including under an explicit gender-attribution prompt that asks models to report whether gender influenced their choice. Cross-gender role swaps reveal a consistent female-proposed-decision asymmetry, while models often attribute flipped decisions to No Influence or other non-gender factors. Further analysis shows that gender effects concentrate near less determinate value boundaries and under more severe decision contexts, suggesting that gender cues act as local boundary-shifting factors rather than global overrides of value reasoning. Value rankings remain largely stable, but ordered value-pair trade-offs shift unevenly across role-gender configurations. These results show that gender can enter LLM value trade-offs behaviorally while remaining obscured in self-attribution, motivating controlled behavioral audits beyond explanation-based evaluation.

大模型性别偏见决策公平性行为审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。