用博弈论让AI在跨文化对话中达成更公平的共识
A Game-Theoretic Negotiation Framework for Cross-Cultural Consensus in LLMs
- 将文化共识建模为纳什均衡,用博弈算法模拟跨文化谈判
- 在多文化数据上验证,生成的共识更平衡且一致性更高
- 适合关注AI公平性、文化多样性研究者使用
大型语言模型正影响全球价值体系,但其常因忽视少数群体价值而呈现明显的WEIRD(西方、受教育、工业化、富裕、民主)文化偏见。这种单一文化视角可能强化主导价值观,边缘化多元文化观点,阻碍公平包容AI的发展。本文提出一种系统性框架,旨在提升LLMs间的跨文化共识质量。将共识建模为纳什均衡,采用基于策略空间响应预言机(PSRO)的博弈论谈判方法,模拟有组织的跨文化协商过程。通过世界价值观调查(WVS)数据构建区域文化代理进行评估,并引入两个新量化指标:基于困惑度的接受度与价值观自一致性,以衡量共识效果。实验表明,该方法生成的共识质量更高,妥协更均衡,相比基线显著缓解了WEIRD偏见,通过公平渐进的谈判步骤引导代理达成收敛。
原文摘要 · Abstract (English)
The increasing prevalence of large language models (LLMs) is influencing global value systems. However, these models frequently exhibit a pronounced WEIRD (Western, Educated, Industrialized, Rich, Democratic) cultural bias due to lack of attention to minority values. This monocultural perspective may reinforce dominant values and marginalize diverse cultural viewpoints, posing challenges for the development of equitable and inclusive AI systems. In this work, we introduce a systematic framework designed to boost fair and robust cross-cultural consensus among LLMs. We model consensus as a Nash Equilibrium and employ a game-theoretic negotiation method based on Policy-Space Response Oracles (PSRO) to simulate an organized cross-cultural negotiation process. To evaluate this approach, we construct regional cultural agents using data transformed from the World Values Survey (WVS). Beyond the conventional model-level evaluation method, We further propose two quantitative metrics, Perplexity-based Acceptence and Values Self-Consistency, to assess consensus outcomes. Experimental results indicate that our approach generates consensus of higher quality while ensuring more balanced compromise compared to baselines. Overall, it mitigates WEIRD bias by guiding agents toward convergence through fair and gradual negotiation steps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。