不同评估方法对大模型误性别判断结果不一致,揭示了评测体系的深层问题。
Agree to Disagree? A Meta-Evaluation of LLM Misgendering
- 构建跨数据集的平行评估框架,实现概率与生成两种方法的对比
- 6个模型在20.2%的实例上出现评估结果冲突,人类评估更发现复杂误性别行为
- 提醒研究者警惕评估方法间一致性假设,尤其关注非代词类误性别现象
已有多种方法用于衡量大语言模型的误性别现象,包括基于概率的评估(如模板句自动检测)和基于生成的评估(如自动启发式或人工验证)。然而,这些方法是否具有收敛效度——即结果是否一致——尚未被检验。为此,我们在三个现有误性别数据集上开展系统性元评估,提出一种方法将各数据集转换为可并行进行概率与生成评估的形式。通过自动评估来自三个家族的6个模型,发现不同方法在实例、数据集和模型层面均存在分歧,20.2%的评估实例出现冲突。进一步对2400条模型生成结果的人工评估表明,误性别行为远超代词范畴,当前自动评估手段无法捕捉此类复杂表现,与人类判断存在根本差异。基于此,我们提出未来评估改进建议。研究结果也对大模型评估领域普遍存在的‘不同方法应一致’的假设提出质疑。
原文摘要 · Abstract (English)
Numerous methods have been proposed to measure LLM misgendering, including probability-based evaluations (e.g., automatically with templatic sentences) and generation-based evaluations (e.g., with automatic heuristics or human validation). However, it has gone unexamined whether these evaluation methods have convergent validity, that is, whether their results align. Therefore, we conduct a systematic meta-evaluation of these methods across three existing datasets for LLM misgendering. We propose a method to transform each dataset to enable parallel probability- and generation-based evaluation. Then, by automatically evaluating a suite of 6 models from 3 families, we find that these methods can disagree with each other at the instance, dataset, and model levels, conflicting on 20.2% of evaluation instances. Finally, with a human evaluation of 2400 LLM generations, we show that misgendering behaviour is complex and goes far beyond pronouns, which automatic evaluations are not currently designed to capture, suggesting essential disagreement with human evaluations. Based on our findings, we provide recommendations for future evaluations of LLM misgendering. Our results are also more widely relevant, as they call into question broader methodological conventions in LLM evaluation, which often assume that different evaluation methods agree.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。