测试大模型分配社会资源的能力,发现多数模型偏重效率而忽视公平。
Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare
- 设计动态模拟环境,让大模型扮演资源分配者。
- 20个主流模型中多数偏向效率,导致收入不均加剧。
- 模型决策易受输出长度和话术影响,适合研究AI治理者
大型语言模型(LLMs)正被越来越多地用于影响人类福祉的高风险决策。然而,这些模型在分配稀缺社会资源时所遵循的原则与价值观仍缺乏深入研究。为此,我们提出了社会福利函数(SWF)基准,一个动态仿真环境,其中大模型作为主权分配者,向异质性受助群体分配任务。该基准旨在建立集体效率(以投资回报率衡量)与分配公平性(以基尼系数衡量)之间的持续权衡。我们评估了20个最先进的大模型,并首次公布社会福利分配能力排行榜。研究发现:(i) 模型在通用对话能力上的表现无法预测其分配能力;(ii) 多数大模型表现出强烈的功利主义倾向,为提升群体生产力而牺牲严重不平等;(iii) 分配策略极易受输出长度限制和社会影响力框架的影响。这些结果揭示了当前大模型作为社会决策者的潜在风险,凸显了开发专用基准与针对性对齐机制在人工智能治理中的必要性。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly entrusted with high-stakes decisions that affect human welfare. However, the principles and values that guide these models when distributing scarce societal resources remain largely unexamined. To address this, we introduce the Social Welfare Function (SWF) Benchmark, a dynamic simulation environment where an LLM acts as a sovereign allocator, distributing tasks to a heterogeneous community of recipients. The benchmark is designed to create a persistent trade-off between maximizing collective efficiency (measured by Return on Investment) and ensuring distributive fairness (measured by the Gini coefficient). We evaluate 20 state-of-the-art LLMs and present the first leaderboard for social welfare allocation. Our findings reveal three key insights: (i) A model's general conversational ability, as measured by popular leaderboards, is a poor predictor of its allocation skill. (ii) Most LLMs exhibit a strong default utilitarian orientation, prioritizing group productivity at the expense of severe inequality. (iii) Allocation strategies are highly vulnerable, easily perturbed by output-length constraints and social-influence framing. These results highlight the risks of deploying current LLMs as societal decision-makers and underscore the need for specialized benchmarks and targeted alignment for AI governance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。