arXiv:2507.16557cs.CLcs.LG2025-07中稿 · the 6th Workshop o…被引 4

针对德语大模型性别偏见,构建5个评估数据集并揭示语言特异性挑战

Exploring Gender Bias in Large Language Models: An In-depth Dive into the German Language

  • 基于德语特性设计5个性别偏见评估数据集,覆盖多种测量方法
  • 在8个双语大模型上发现男性职业术语歧义及中性词影响性别认知
  • 强调需为不同语言定制评估框架,适合多语言偏见研究者

近年来,已有多种方法被提出用于评估大语言模型(LLMs)中的性别偏见。一个关键挑战在于,最初为英语设计的偏见测量方法在应用于其他语言时的可迁移性问题。本文旨在推动该研究方向,提出5个用于德语大语言模型性别偏见评估的数据集。这些数据集基于成熟的性别偏见概念构建,可通过多种方法访问。对8个双语大模型的分析结果揭示了德语中独特的偏见挑战,包括男性职业术语的模糊解读以及看似中性的名词对性别感知的影响。本研究有助于理解跨语言背景下大模型的性别偏见,并强调了开发定制化评估框架的必要性。

原文摘要 · Abstract (English)

In recent years, various methods have been proposed to evaluate gender bias in large language models (LLMs). A key challenge lies in the transferability of bias measurement methods initially developed for the English language when applied to other languages. This work aims to contribute to this research strand by presenting five German datasets for gender bias evaluation in LLMs. The datasets are grounded in well-established concepts of gender bias and are accessible through multiple methodologies. Our findings, reported for eight multilingual LLM models, reveal unique challenges associated with gender bias in German, including the ambiguous interpretation of male occupational terms and the influence of seemingly neutral nouns on gender perception. This work contributes to the understanding of gender bias in LLMs across languages and underscores the necessity for tailored evaluation frameworks.

性别偏见大模型德语评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。