警惕盲目模仿人类行为的AI对齐,可能引入统计偏差。
A Statistical Case Against Empirical Human-AI Alignment
- 反对直接复制人类行为数据进行对齐
- 提出事后验证与原则性对齐作为替代方案
- 适合关注AI伦理与可靠性的研究者
经验式人机对齐旨在让AI系统的行为符合观察到的人类行为。尽管目标可贵,但我们认为这种做法可能无意中引入统计偏差,需谨慎对待。本文主张避免简单的经验对齐,转而推荐基于原则的对齐与事后经验对齐作为替代方案。通过语言模型的人类中心解码等具体例子,论证了该观点的合理性。
原文摘要 · Abstract (English)
Empirical human-AI alignment aims to make AI systems act in line with observed human behavior. While noble in its goals, we argue that empirical alignment can inadvertently introduce statistical biases that warrant caution. This position paper thus advocates against naive empirical alignment, offering prescriptive alignment and a posteriori empirical alignment as alternatives. We substantiate our principled argument by tangible examples like human-centric decoding of language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。