修复性别偏见评估数据集缺陷,提出新基准测试模型对代词的分辨能力。
WinoPron: Revisiting English Winogender Schemas for Consistency, Coverage, and Grammatical Case
- 重构原始数据,修正代词等价性、模板违规等问题,提升评测可靠性。
- 发现所有模型在宾格代词上表现更差,尤其在复杂句式中准确率下降显著。
- 提出多维度偏见评估方法,区分同组代词不同形式的偏差差异。
尽管衡量共指消解中的偏见与鲁棒性至关重要,但这些评估的有效性取决于所用工具的质量。Winogender Schemas(Rudinger et al., 2018)是用于评估共指消解中性别偏见的有影响力的语料库,但深入分析发现其存在若干问题:将不同代词形式视为等价、违反模板约束、拼写错误等,削弱了其作为可靠评估工具的价值。本文识别并修正这些问题,构建新数据集 WinoPron。利用该数据集,我们评估了两种前沿监督式共指消解系统——SpanBERT 以及五种规模的 FLAN-T5,结果表明所有模型在宾格代词(如 him, her)上的解析准确率均显著更低。此外,本文提出一种超越二元判断的代词偏见评估方法,发现偏见特征不仅存在于代词类别之间(如 he vs. she),也存在于同一类别内的不同表面形式之间(如 him vs. his)。
原文摘要 · Abstract (English)
While measuring bias and robustness in coreference resolution are important goals, such measurements are only as good as the tools we use to measure them. Winogender Schemas (Rudinger et al., 2018) are an influential dataset proposed to evaluate gender bias in coreference resolution, but a closer look reveals issues with the data that compromise its use for reliable evaluation, including treating different pronominal forms as equivalent, violations of template constraints, and typographical errors. We identify these issues and fix them, contributing a new dataset: WinoPron. Using WinoPron, we evaluate two state-of-the-art supervised coreference resolution systems, SpanBERT, and five sizes of FLAN-T5, and demonstrate that accusative pronouns are harder to resolve for all models. We also propose a new method to evaluate pronominal bias in coreference resolution that goes beyond the binary. With this method, we also show that bias characteristics vary not just across pronoun sets (e.g., he vs. she), but also across surface forms of those sets (e.g., him vs. his).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。