多智能体协作中,公平性是交互产生的结果,而非单个模型的属性。
Beyond Arrow's Impossibility: Fairness as an Emergent Property of Multi-Agent Collaboration
- 通过两智能体辩论框架,研究公平性在交互中的涌现机制。
- 单独代理无法满足公平标准,但联合决策可达成单一代理无法实现的公平结果。
- 适用于研究去中心化系统中的公平性设计,尤其关注多智能体协同场景。
语言模型的公平性通常被视为单一中心优化模型的属性。随着大语言模型日益具备自主性,我们提出公平性可通过交互与协作产生。本文通过受控的医院分诊框架,让两个智能体进行三轮结构化辩论:一方通过检索增强生成(RAG)对齐特定伦理框架,另一方则未对齐或被恶意提示以偏好特定人口群体而非临床需求。研究发现,对齐显著影响谈判策略与资源分配模式;单独来看,双方分配均不满足伦理要求,但其最终联合分配却能符合公平标准,这是任一智能体单独行动无法达到的。对齐智能体通过博弈而非强制纠正偏差,作为校正模块恢复边缘群体的资源获取,而不完全改变有偏对手。此外,即使明确对齐的智能体也表现出对某些框架的内在倾向,符合大模型普遍存在的左倾趋势。这一局限与阿罗不可能定理相关:任何聚合机制都无法同时满足集体理性的所有理想条件,多智能体协商只是绕过而非解决该约束。研究将公平性重新定义为去中心化交互过程中的涌现性、程序性属性,系统的整体表现才是评估公平性的合适单位。
原文摘要 · Abstract (English)
Fairness in language models is typically studied as a property of a single, centrally optimized model. As large language models become increasingly agentic, we propose that fairness emerges through interaction and exchange. We study this via a controlled hospital triage framework in which two agents negotiate over three structured debate rounds. One agent is aligned to a specific ethical framework via retrieval-augmented generation (RAG), while the other is either unaligned or adversarially prompted to favor demographic groups over clinical need. We find that alignment systematically shapes negotiation strategies and allocation patterns, and that neither agent's allocation is ethically adequate in isolation, yet their joint final allocation can satisfy fairness criteria that neither would have reached alone. Aligned agents partially moderate bias through contestation rather than override, acting as corrective patches that restore access for marginalized groups without fully converting a biased counterpart. We further observe that even explicitly aligned agents exhibit intrinsic biases toward certain frameworks, consistent with known left-leaning tendencies in LLMs. We connect these limits to Arrow's Impossibility Theorem: no aggregation mechanism can simultaneously satisfy all desiderata of collective rationality, and multi-agent deliberation navigates rather than resolves this constraint. Our results reposition fairness as an emergent, procedural property of decentralized agent interaction, and the system rather than the individual agent as the appropriate unit of evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。