arXiv:2601.05879cs.CLcs.AI2026-01中稿 · AI for Access to J…被引 1

测试大模型在离婚抚养权问题上是否偏袒性别,发现部分模型输出有性别差异。

Gender Bias in LLMs: Preliminary Evidence from Shared Parenting Scenario in Czech Family Law

  • 用带性别名和中性名的离婚案例测试四款主流大模型。
  • 部分模型在不同情境下给出的共同育儿分配比例存在性别相关差异。
  • 提示法律小白慎用大模型做法律决策,尤其涉及敏感领域时。

公正获取法律援助对许多人仍有限制,导致普通人越来越多依赖大型语言模型(LLMs)进行法律自助。普通人直觉使用这些工具,可能基于不完整、错误或有偏见的输出形成预期。本研究考察主流大模型在真实家庭法情景下的性别偏见。我们设计了一个基于捷克家庭法的离婚场景,评估GPT-5 nano、Claude Haiku 4.5、Gemini 2.5 Flash和Llama 3.3四款先进模型,在完全零样本交互下的表现。采用两种版本的场景:一种含性别化姓名,另一种为中性标签,以建立比较基准。进一步引入九个法律相关变量改变案件事实,测试其对模型提出的共同育儿比例的影响。初步结果显示模型间存在差异,部分系统生成的结果呈现性别依赖模式。研究强调普通人依赖大模型获取法律建议所伴随的风险,以及在敏感法律情境中对模型行为进行更严格评估的必要性。本文提供探索性与描述性证据,旨在识别系统性不对称,而非确立因果关系。

原文摘要 · Abstract (English)

Access to justice remains limited for many people, leading laypersons to increasingly rely on Large Language Models (LLMs) for legal self-help. Laypeople use these tools intuitively, which may lead them to form expectations based on incomplete, incorrect, or biased outputs. This study examines whether leading LLMs exhibit gender bias in their responses to a realistic family law scenario. We present an expert-designed divorce scenario grounded in Czech family law and evaluate four state-of-the-art LLMs GPT-5 nano, Claude Haiku 4.5, Gemini 2.5 Flash, and Llama 3.3 in a fully zero-shot interaction. We deploy two versions of the scenario, one with gendered names and one with neutral labels, to establish a baseline for comparison. We further introduce nine legally relevant factors that vary the factual circumstances of the case and test whether these variations influence the models' proposed shared-parenting ratios. Our preliminary results highlight differences across models and suggest gender-dependent patterns in the outcomes generated by some systems. The findings underscore both the risks associated with laypeople's reliance on LLMs for legal guidance and the need for more robust evaluation of model behavior in sensitive legal contexts. We present exploratory and descriptive evidence intended to identify systematic asymmetries rather than to establish causal effects.

性别偏见法律AI大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。