用博弈论方法生成更公平的共识文本,让每个人声音都得到更好体现。
Generating Fair Consensus Statements with Social Choice on Token-Level MDPs
- 将文本生成建模为多目标令牌级马尔可夫决策过程,每条意见对应一个目标。
- 搜索算法优化平等福利,使最不满意者对齐度提升,优于基线方法。
- 适用于需要多方共识的场景,如政策讨论、多元观点整合。
当前基于大语言模型的共识文本生成框架缺乏内在结构,无法提供可证明的公平性保障。本文将任务建模为多目标、令牌级马尔可夫决策过程(MDP),每个目标对应一个代理的偏好。各代理的令牌级奖励由其策略(如个性化语言模型)推导得出,利用策略隐含定义最优Q函数的特性,在无需显式价值函数的情况下量化每步生成的奖励(Rafailov et al., 2024)。该MDP框架可借助社会选择理论进行形式化分析。提出两种基于社会选择理论的方法:一是构造随机生成策略,保证事前核心稳定性,源自最大化比例公平(纳什福利)的完整陈述分布;二是针对生成单一陈述,采用搜索算法在MDP框架内最大化平等福利。实验表明,使用语言模型作为代理策略,基于平等目标的搜索生成的共识文本,在最差情况下的代理对齐度显著优于基线方法,包括Habermas Machine(Tessler et al., 2024)。
原文摘要 · Abstract (English)
Current frameworks for consensus statement generation with large language models lack the inherent structure needed to provide provable fairness guarantees when aggregating diverse free-form opinions. We model the task as a multi-objective, token-level Markov Decision Process (MDP), where each objective corresponds to an agent's preference. Token-level rewards for each agent are derived from their policy (e.g., a personalized language model). This approach utilizes the finding that such policies implicitly define optimal Q-functions, providing a principled way to quantify rewards at each generation step without a value function (Rafailov et al., 2024). This MDP formulation creates a formal structure amenable to analysis using principles from social choice theory. We propose two approaches grounded in social choice theory. First, we propose a stochastic generation policy guaranteed to be in the ex-ante core, extending core stability concepts from voting theory to text generation. This policy is derived from an underlying distribution over complete statements that maximizes proportional fairness (Nash Welfare). Second, for generating a single statement, we target the maximization of egalitarian welfare using search algorithms within the MDP framework. Empirically, experiments using language models to instantiate agent policies show that search guided by the egalitarian objective generates consensus statements with improved worst-case agent alignment compared to baseline methods, including the Habermas Machine (Tessler et al., 2024).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。