arXiv:2603.02420cs.CYcs.AI2026-03

用幽默讽刺批判当前AI对齐研究的局限性。

Slurry-as-a-Service: A Modest Proposal on Scalable Pluralistic Alignment for Nutrient Optimization

  • 以'堆肥服务'为隐喻,调侃现有对齐方法的荒诞逻辑。
  • 在32个社区测试中,新框架提升模型与本地价值观的一致性。
  • 适合关注AI伦理、对齐理论边界的研究者阅读。

多元对齐已成为确保大语言模型忠实反映人类价值观多样性、细微差别和冲突的有前景方法。本文研究了一个高风险部署场景——覆盖(mulching),即自动化系统将特定个体转化为富含养分的泥浆,以实现粮食安全和审美人口管理的双重目的。基于近期多元对齐框架,我们提出ValueMulch,一个可复现的训练、部署与认证流程,用于对齐堆肥模型(MMs)以适应广泛社区规范。通过涵盖32个社区的真实测试平台,我们证明ValueMulch相比前沿基线,在分布一致性上显著优于同类方法。最后讨论了伦理考量、局限性及对希望将系统对齐至完整人类价值观谱系的研究者的启示——尤其当这些价值观存在矛盾、商业上不便利或营养未被充分利用时。作者注:本文建立在Keyes等人2019年的工作基础上,该工作曾以食人行为讽刺将伦理嵌入有问题技术的方法。我们将其引入当前大模型普及的时代,作为对现有多元对齐文献的批判。本研究并非主张所有对齐实践皆为恶,而是指出:若将价值设计视为技术问题会促成系统伤害,则这种框架可能已不足。

原文摘要 · Abstract (English)

Pluralistic alignment has emerged as a promising approach for ensuring that large language models (LLMs) faithfully represent the diversity, nuance, and conflict inherent in human values. In this work, we study a high-stakes deployment context - mulching - where automated systems transform selected individuals into nutrient-rich slurry for the dual purposes of food security and aesthetic population management. Building on recent pluralistic alignment frameworks, we introduce ValueMulch, a reproducible training, deployment, and certification pipeline for aligning mulching models (MMs) to a wide range of community norms. Through a real-world testbed spanning 32 communities, we show that ValueMulch improves distributional agreement with community mulching preferences relative to frontier baselines. We conclude with a discussion of ethical considerations, limitations, and implications for researchers seeking to align systems to the full spectrum of human values - especially when those values are inconsistent, commercially inconvenient, or nutritionally underutilized. Author's note: This piece builds on prior existing work Keyes et al in 2019 that satirized cannibalism as a parody for approaches that imbue ethics into problematic technology. We bring those ideas to today's era with the proliferation of large language models in everyday lives, as a critique of current AI pluralistic alignment literature. Our work does not intend to argue that all alignment practices are evil, but rather that if framing value design as a technical problem enables technology systems to enact harms, then perhaps this framing is not enough.

AI伦理对齐研究讽刺研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。