用学习理论解决大规模集体决策中的代表性问题,为AI对齐提供新思路。
Representative Social Choice: From Learning Theory to AI Alignment
- 将代表制社会选择建模为统计学习问题,通过样本推断整体偏好。
- 证明了社会选择机制的泛化能力,关键在样本规模与误差关系。
- 适用于语言模型对齐、立法等大规模民主决策场景。
社会选择理论研究群体偏好聚合,广泛应用于人类机制设计及语言模型的民主对齐。本文提出代表制社会选择框架,用于处理个体与议题数量过大而无法直接考虑所有偏好的情形,如陪审团审判、立法、公司治理及近年的语言模型对齐。该框架基于有限的个体-议题样本进行社会选择决策。我们证明,许多代表制社会选择的核心问题可转化为统计学习问题,并利用机器学习理论分析社会选择机制的泛化性能。进一步,我们构建了代表制社会选择的公理体系,借助新的组合分析工具,证明了类阿罗不可能性定理。本框架首次将代表方法引入社会选择领域,开辟了社会选择、学习理论与人工智能对齐交叉研究的新方向。
原文摘要 · Abstract (English)
Social choice theory is the study of preference aggregation across a population, used both in mechanism design for human agents and in the democratic alignment of language models. In this study, we propose the representative social choice framework for the modeling of democratic representation in collective decisions, where the number of issues and individuals are too large for mechanisms to consider all preferences directly. These scenarios are widespread in real-world decision-making processes, such as jury trials, legislation, corporate governance, and, more recently, language model alignment. In representative social choice, the population is represented by a finite sample of individual-issue pairs based on which social choice decisions are made. We show that many of the deepest questions in representative social choice can be formulated as statistical learning problems, and prove the generalization properties of social choice mechanisms using the theory of machine learning. We further formulate axioms for representative social choice, and prove Arrow-like impossibility theorems with new combinatorial tools of analysis. Our framework introduces the representative approach to social choice, opening up research directions at the intersection of social choice, learning theory, and AI alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。