评测大模型在三种语言中性别感知形态生成的能力
MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological Generation

- 构建跨语言的性别敏感形态生成评测集
- 15个主流多语言模型在性别转换任务上表现参差
- 为包容性自然语言处理提供诊断工具
尽管多语言大模型在翻译和问答等高层任务上表现良好,其对语法性别和形态一致性的处理能力仍缺乏深入研究。在形态丰富的语言中,性别影响动词变位、代词使用,甚至第一人称表达中显式或隐式的性别指涉。我们提出MORPHOGEN,一个大规模、基于形态学的基准数据集,用于评估法语、阿拉伯语和印地语这三种类型多样且具有语法性别的语言中的性别感知生成能力。核心任务GENFORM要求模型将第一人称句子改写为相反性别,同时保持语义和结构不变。我们构建了一个高质量的合成数据集,并对15个主流多语言大模型(2B-70B参数量)进行了基准测试。结果揭示了当前模型在形态性别处理上的显著差距与有趣现象。MORPHOGEN为性别敏感语言建模提供了聚焦的诊断视角,也为未来包容性与形态敏感的NLP研究奠定基础。
原文摘要 · Abstract (English)
While multilingual large language models (LLMs) perform well on high-level tasks like translation and question answering, their ability to handle grammatical gender and morphological agreement remains underexplored. In morphologically rich languages, gender influences verb conjugation, pronouns, and even first-person constructions with explicit and implicit mentions of gender. We introduce MORPHOGEN, a morphologically grounded large-scale benchmark dataset for evaluating gender-aware generation in three typologically diverse grammatically gendered languages: French, Arabic, and Hindi. The core task, GENFORM, requires models to rewrite a first-person sentence in the opposite gender while preserving its meaning and structure. We construct a high-quality synthetic dataset spanning these three languages and benchmark 15 popular multilingual LLMs (2B-70B) on their ability to perform this transformation. Our results reveal significant gaps and interesting insights into how current models handle morphological gender. MORPHOGEN provides a focused diagnostic lens for gender-aware language modeling and lays the groundwork for future research on inclusive and morphology-sensitive NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。