arXiv:2504.02973cs.CL2025-04

用概率模型模拟代词学习,让系统更尊重跨性别者多样表达

A Bayesian account of pronoun and neopronoun acquisition

  • 基于嵌套中式餐厅特许经营模型,建模个体差异的代词选择
  • 能解释不同人对代词/新式代词的接受速度差异
  • 适合做包容性语言系统的开发者参考

跨性别群体在使用自选指称(如姓名或新式代词)时面临显著公平挑战。说话者常将误用归因于‘无意’,声称这是长期主流语言习惯固化及外貌与指称之间刻板预期所致。本文主张显式建模个体在代词选择上的差异,提出基于嵌套中式餐厅特许经营过程(nCRFP)的概率图模型,用于解释灵活指称(如自选名、新式代词)现象。该模型不依赖词汇共现统计,突破传统语言模型的形式-意义映射局限。结果表明,该模型可有效捕捉代词或名字在符号知识中整合速度的个体差异,并使计算系统在保持灵活性的同时,尊重性别表达多元的跨性别者。

原文摘要 · Abstract (English)

A major challenge to equity among members of queer communities is the use of one's chosen forms of reference, such as personal names or pronouns. Speakers often dismiss their misuses of pronouns as "unintentional", and claim that their errors reflect many decades of fossilized mainstream language use, as well as attitudes or expectations about the relationship between one's appearance and acceptable forms of reference. We argue for explicitly modeling individual differences in pronoun selection and present a probabilistic graphical modeling approach based on the nested Chinese Restaurant Franchise Process (nCRFP) (Ahmed et al., 2013) to account for flexible pronominal reference such as chosen names and neopronouns while moving beyond form-to-meaning mappings and without lexical co-occurrence statistics to learn referring expressions, as in contemporary language models. We show that such a model can account for variability in how quickly pronouns or names are integrated into symbolic knowledge and can empower computational systems to be both flexible and respectful of queer people with diverse gender expression.

代词模型性别表达概率建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。