用幽默讽刺批判当前AI对齐研究的局限性。
Slurry-as-a-Service: A Modest Proposal on Scalable Pluralistic Alignment for Nutrient Optimization
- 以'堆肥服务'为隐喻,调侃现有对齐方法的荒诞逻辑。
- 在32个社区测试中,新框架提升模型与本地价值观的一致性。
- 适合关注AI伦理、对齐理论边界的研究者阅读。
多元对齐已成为确保大语言模型忠实反映人类价值观多样性、细微差别和冲突的有前景方法。本文研究了一个高风险部署场景——覆盖(mulching),即自动化系统将特定个体转化为富含养分的泥浆,以实现粮食安全和审美人口管理的双重目的。基于近期多元对齐框架,我们提出ValueMulch,一个可复现的训练、部署与认证流程,用于对齐堆肥模型(MMs)以适应广泛社区规范。通过涵盖32个社区的真实测试平台,我们证明ValueMulch相比前沿基线,在分布一致性上显著优于同类方法。最后讨论了伦理考量、局限性及对希望将系统对齐至完整人类价值观谱系的研究者的启示——尤其当这些价值观存在矛盾、商业上不便利或营养未被充分利用时。作者注:本文建立在Keyes等人2019年的工作基础上,该工作曾以食人行为讽刺将伦理嵌入有问题技术的方法。我们将其引入当前大模型普及的时代,作为对现有多元对齐文献的批判。本研究并非主张所有对齐实践皆为恶,而是指出:若将价值设计视为技术问题会促成系统伤害,则这种框架可能已不足。
原文摘要 · Abstract (English)
Pluralistic alignment has emerged as a promising approach for ensuring that large language models (LLMs) faithfully represent the diversity, nuance, and conflict inherent in human values. In this work, we study a high-stakes deployment context - mulching - where automated systems transform selected individuals into nutrient-rich slurry for the dual purposes of food security and aesthetic population management. Building on recent pluralistic alignment frameworks, we introduce ValueMulch, a reproducible training, deployment, and certification pipeline for aligning mulching models (MMs) to a wide range of community norms. Through a real-world testbed spanning 32 communities, we show that ValueMulch improves distributional agreement with community mulching preferences relative to frontier baselines. We conclude with a discussion of ethical considerations, limitations, and implications for researchers seeking to align systems to the full spectrum of human values - especially when those values are inconsistent, commercially inconvenient, or nutritionally underutilized. Author's note: This piece builds on prior existing work Keyes et al in 2019 that satirized cannibalism as a parody for approaches that imbue ethics into problematic technology. We bring those ideas to today's era with the proliferation of large language models in everyday lives, as a critique of current AI pluralistic alignment literature. Our work does not intend to argue that all alignment practices are evil, but rather that if framing value design as a technical problem enables technology systems to enact harms, then perhaps this framing is not enough.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。