让AI随社会价值观变化动态调整,避免固守旧标准。
Adaptive Pluralistic Alignment: A pipeline for dynamic artificial democracy

- 用低秩分解构建个性化奖励模型,高效捕捉多元偏好。
- 通过投票机制综合多个模型输出,实现动态决策。
- 无需重训即可更新系统,适合需要持续演进的AI应用。
现有对齐方法针对固定偏好,随社会规范演变易导致价值固化。本文提出自适应多元对齐(APA),一种模块化流水线,可在不重复昂贵预训练或大规模数据收集的前提下,使对齐后的AI系统动态追踪价值变迁、避免价值固化。APA包含三个阶段:(1) 通过低秩奖励基分解学习紧凑的个性化奖励模型;(2) 利用这些模型作为评审团,基于社会选择理论投票从候选输出中选出最优解;(3) 当价值变迁时,仅需拟合新的标注者权重,而保留固定的奖励基以实现高效更新。该系统兼具高效性、可解释性、可控性和模块化特性。我们在PRISM多用户对齐数据集上实现概念验证,并使用模拟历史标注者进行初步分析,结果表明评审团构成与投票规则显著影响结果,尤其在评审偏好异质时更为明显。完整代码与生成的偏好数据集已公开于https://github.com/RachelFreedman/apa。
原文摘要 · Abstract (English)
Prevailing alignment methods target a fixed set of preferences and therefore risk forcing value lock-in as societal norms evolve over time. We introduce Adaptive Pluralistic Alignment (APA), a modular pipeline for updating pluralistically aligned AI systems to track evolving values and avoid value lock-in without repeating costly pretraining or large-scale data collection. APA has three stages: (1) learning compact personalized reward models via low-rank reward basis decomposition, (2) using these models as a jury that collectively selects among candidate outputs through social-choice-theoretic voting, and (3) efficiently adapting the jury over time by fitting new annotator weights over the fixed reward bases as values shift. The resulting system is efficient, explainable, steerable, and modular. We implement a proof-of-concept instantiation using the PRISM multi-user alignment dataset and simulated historical annotators, and provide preliminary analysis showing that jury composition and the choice of voting rule can substantially affect outcomes, particularly when jury preferences are heterogeneous. We provide full code and resulting preference datasets at https://github.com/RachelFreedman/apa.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。