42种语言的模型生成中减少误指性别,兼顾文化适配性。
A Multilingual, Culture-First Approach to Addressing Misgendering in LLM Applications
- 基于参与式设计构建多语言性别指代防护机制。
- 在会议摘要任务中降低所有语言的误指率且不损失质量。
- 适合关注跨文化AI伦理与包容性设计的研究者。
误指性别是指用不符合个体自选身份的性别称谓称呼他人,会造成严重伤害。英语已有明确规避策略,如使用‘they’代词,但其他语言因语法与文化差异面临独特挑战。本文针对42种语言和方言,提出评估与缓解误指性别的方法,采用参与式设计开发适用于各语言的防护机制。在标准大模型应用(会议记录摘要)中,通过人机协同的数据生成与标注流程验证效果。结果表明,所提防护机制显著降低各类语言生成摘要中的误指率,且未造成质量下降。该人机协同方法为多语言多文化环境下可扩展的负责任AI方案提供了可行路径。研究发布涵盖42种语言的防护机制、合成数据集及人工与大模型评分,以促进该领域深入研究。
原文摘要 · Abstract (English)
Misgendering is the act of referring to someone by a gender that does not match their chosen identity. It marginalizes and undermines a person's sense of self, causing significant harm. English-based approaches have clear-cut approaches to avoiding misgendering, such as the use of the pronoun ``they''. However, other languages pose unique challenges due to both grammatical and cultural constructs. In this work we develop methodologies to assess and mitigate misgendering across 42 languages and dialects using a participatory-design approach to design effective and appropriate guardrails across all languages. We test these guardrails in a standard LLM-based application (meeting transcript summarization), where both the data generation and the annotation steps followed a human-in-the-loop approach. We find that the proposed guardrails are very effective in reducing misgendering rates across all languages in the summaries generated, and without incurring loss of quality. Our human-in-the-loop approach demonstrates a method to feasibly scale inclusive and responsible AI-based solutions across multiple languages and cultures. We release the guardrails and synthetic dataset encompassing 42 languages, along with human and LLM-judge evaluations, to encourage further research on this subject.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。