让多个AI代理学会不同性别歧视判断视角,更真实地模拟人类分歧。
Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization
- 用聚类方法分离标注者视角,为每类训练专属AI代理。
- 团队级奖励机制让代理精准保持各自群体的判断风格。
- 适用于需要理解多元社会观点的伦理审查与公平性研究。
当人们标注文本中的性别歧视时,常存在分歧,这并非因某方错误,而是感知差异所致。现有NLP系统通常通过多数投票忽略这种分歧。我们提出多智能体视角偏好优化(MAP-PO)框架,保留不同视角。在英文和西班牙语的EXIST 2024微博数据集上,首先根据标注行为而非人口属性聚类标注者;然后为每类训练一个大语言模型代理,使其复现该类标注行为,并通过结合个体与团队奖励的偏好优化进行协调。在两种语言与两种主干模型共四种设置下评估:每个代理是否复现所属集群的标注,以及代理群体是否复现多数标签。四个设置中均发现:未微调时代理表现几乎相同,说明必须进行聚类特异训练;仅使用本集群标签训练会使代理偏离其所属集群,而引入共享团队级训练信号可稳定保持各代理对所属集群的校准性。
原文摘要 · Abstract (English)
When people label text for sexism, they often disagree, and not because some of them are wrong: they genuinely perceive sexism differently. Most NLP systems discard this disagreement by collapsing it into a majority vote. We propose the Multi-Agent Perspectivist Preference Optimization (MAP-PO) framework to keep these different perspectives. On the EXIST 2024 dataset of labeled English and Spanish tweets, we first cluster annotators by their labeling behavior rather than their demographic attributes. We then fine-tune one Large Language Model agent per cluster to reproduce that cluster's annotation behavior, and coordinate the agents with preference optimization that combines individual and team-level rewards. We evaluate MAP-PO in four settings defined by two languages and two backbone language models, asking whether each agent reproduces the annotations of its own cluster and whether the agents together reproduce the majority label. Two findings hold in all four settings. First, without fine-tuning the agents behave almost identically, so cluster-specific training is necessary. Second, we show that training each agent only on the labels of its own cluster pushes the agents far beyond the clusters they should represent, while adding a shared team-level training signal consistently keeps each agent calibrated to its cluster.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。