测试大模型在资源分配中的公平性,发现其与人类偏好严重不符。
Distributive Fairness in Large Language Models: Evaluating Alignment with Human Values
- 让大模型从预设选项中选方案,比自由生成更符合公平原则
- 多数模型无法用金钱调节不平等,且对公平概念理解不足
- 适合关注AI伦理、政策设计的研究者参考
随着大语言模型(LLMs)在社会与经济决策中应用日益广泛,其作为代理行为的公平性引发关注。本文评估了多个主流大模型在资源分配任务中是否符合等价性、无嫉妒性及罗尔斯式最大最小原则等基本公平准则,并考察其与人类偏好的对齐程度。实验表明,当前大模型的回应与人类分布偏好存在显著偏差,且无法有效利用金钱作为可转移资源缓解不平等。然而,当模型被要求从预设选项中选择时,表现明显改善。此外,我们分析了语义因素(如意图或角色设定)和非语义提示变化(如模板或顺序)对响应稳定性的影响。最后,提出若干提升模型与公平原则对齐的潜在策略。
原文摘要 · Abstract (English)
The growing interest in employing large language models (LLMs) for decision-making in social and economic contexts has raised questions about their potential to function as agents in these domains. A significant number of societal problems involve the distribution of resources, where fairness, along with economic efficiency, play a critical role in the desirability of outcomes. In this paper, we examine whether LLM responses adhere to fundamental fairness concepts such as equitability, envy-freeness, and Rawlsian maximin, and investigate their alignment with human preferences. We evaluate the performance of several LLMs, providing a comparative benchmark of their ability to reflect these measures. Our results demonstrate a lack of alignment between current LLM responses and human distributional preferences. Moreover, LLMs are unable to utilize money as a transferable resource to mitigate inequality. Nonetheless, we demonstrate a stark contrast when (some) LLMs are tasked with selecting from a predefined menu of options rather than generating one. In addition, we analyze the robustness of LLM responses to variations in semantic factors (e.g., intentions or personas) or non-semantic prompting changes (e.g., templates or orderings). Finally, we highlight potential strategies aimed at enhancing the alignment of LLM behavior with well-established fairness concepts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。