用波兰语指令微调,让大模型生成更性别平等的文本。
Integrating gender inclusivity into large language models via instruction tuning
- 基于波兰语语法设计包含性提示,引导模型生成中性表达。
- 在多语言和波兰专用模型上验证,显著降低性别偏见输出。
- 适合关注语言公平与社会影响的AI研究者与开发者。
设想一种具有阳性、阴性和中性语法性别区分的语言,但因历史与政治惯例,阳性形式被广泛用于指代男性、女性及混合性别群体——这正是当代波兰语的现实。这种不公平的语言体系导致训练于波兰文本的大语言模型继承并强化了阳性偏见,生成性别失衡的输出。本研究通过使用IPIS数据集(一组人工构建的波兰语性别包容性校对与波兰语-英语翻译指令)对大语言模型进行指令微调,解决该问题。基于理论语言学框架,我们设计了明确包含性指南的系统提示。实验涵盖多语言模型(Llama-8B、Mistral-7B、Mistral-Nemo)和波兰专用模型(Bielik、PLLuM)。该方法旨在将性别包容性作为模型固有特性,提供系统性解决方案,以缓解波兰语生成中的性别偏见。
原文摘要 · Abstract (English)
Imagine a language with masculine, feminine, and neuter grammatical genders, yet, due to historical and political conventions, masculine forms are predominantly used to refer to men, women and mixed-gender groups. This is the reality of contemporary Polish. A social consequence of this unfair linguistic system is that large language models (LLMs) trained on Polish texts inherit and reinforce this masculine bias, generating gender-imbalanced outputs. This study addresses this issue by tuning LLMs using the IPIS dataset, a collection of human-crafted gender-inclusive proofreading in Polish and Polish-to-English translation instructions. Grounded in a theoretical linguistic framework, we design a system prompt with explicit gender-inclusive guidelines for Polish. In our experiments, we IPIS-tune multilingual LLMs (Llama-8B, Mistral-7B and Mistral-Nemo) and Polish-specific LLMs (Bielik and PLLuM). Our approach aims to integrate gender inclusivity as an inherent feature of these models, offering a systematic solution to mitigate gender bias in Polish language generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。