评测大模型对非二元性别代词的处理能力,发现进步有限仍存短板。
Do They Understand Them? An Updated Evaluation on Nonbinary Pronoun Handling in Large Language Models
- 构建新版MISGENDERED+基准,评估五款主流大模型在多种场景下的代词使用
- 二元与中性代词准确率提升,但新式代词和反向推理任务表现不佳
- 适合关注公平性、包容性及大模型社会影响的研究者与开发者
大型语言模型(LLMs)在需兼顾公平与包容性的敏感场景中日益广泛应用。代词使用,尤其是性别中立和新式代词(neopronouns),仍是负责任AI的关键挑战。先前研究如MISGENDERED基准揭示了早期模型在包容性代词处理上的显著缺陷,但受限于过时模型与有限评估。本研究提出MISGENDERED+,一个扩展并更新的基准,用于评估LLMs的代词忠实度。我们对GPT-4o、Claude 4、DeepSeek-V3、Qwen Turbo和Qwen2.5五款代表性模型,在零样本、少样本及性别身份推断任务中进行测试。结果表明相较于以往研究有明显改进,尤其在二元与性别中立代词准确性方面。然而,新式代词及反向推理任务的准确率仍不一致,凸显出在身份敏感推理方面的持续差距。本文讨论其影响、模型特异性观察以及未来包容性AI研究方向。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in sensitive contexts where fairness and inclusivity are critical. Pronoun usage, especially concerning gender-neutral and neopronouns, remains a key challenge for responsible AI. Prior work, such as the MISGENDERED benchmark, revealed significant limitations in earlier LLMs' handling of inclusive pronouns, but was constrained to outdated models and limited evaluations. In this study, we introduce MISGENDERED+, an extended and updated benchmark for evaluating LLMs' pronoun fidelity. We benchmark five representative LLMs, GPT-4o, Claude 4, DeepSeek-V3, Qwen Turbo, and Qwen2.5, across zero-shot, few-shot, and gender identity inference. Our results show notable improvements compared with previous studies, especially in binary and gender-neutral pronoun accuracy. However, accuracy on neopronouns and reverse inference tasks remains inconsistent, underscoring persistent gaps in identity-sensitive reasoning. We discuss implications, model-specific observations, and avenues for future inclusive AI research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。