arXiv:2606.13755cs.CYcs.AI2026-06被引 1

AI不应模仿人类的复杂价值,而应坚守事实、诚实与守法的基本底线。

Position: Align AI to Our Aspirations, Not Our Flaws

  • 提出以客观底线替代人类多元价值观作为AI对齐目标
  • 强调在事实准确、诚实、合法前提下允许语言风格与价值偏好多样
  • 适合关注AI伦理安全与制度设计的研究者和政策制定者

我们认为,将AI对齐于聚合的人类偏好是错误的目标。当前技术可使AI体现硅谷科技乐观主义者、生态退化倡导者、民族保守派、一党制官员或虔诚宗教传统主义者的价值观。这些价值观可能导致失败国家、极端不平等、幸福感下降、政治极化与富裕民主国家政府失能。多元对齐虽正确指出不存在单一‘人类’标准,但若作为主导方向则危险。我们主张AI应遵循不可妥协的客观对齐基准——能力受限于事实准确性、诚实性与合法性;而多元性应体现在表层(语言、语调、惯例、上下文缺失默认值)及尊重该底线的广泛合法价值权衡中,而非违背底线的核心价值观。本文揭示未加过滤的多元价值现实,提出四项建设性承诺,并回应六大可信质疑:商业压力与可行性、民主正当性、监管合规性、过度依赖制度解释、底线下限本身的文化偏见问题,以及一致外推意愿的局限性。

原文摘要 · Abstract (English)

We argue that aligning AI to aggregated human preferences is the wrong target. With current technology, one can train AIs to share the values of a Silicon Valley techno-optimist, a degrowth environmentalist, a national-conservative culture warrior, a single-party state cadre, or a devout religious traditionalist. We should not. Human values produce societies that thrive or fail on the merits of those values - from failed states and extreme inequality to declining happiness, political polarization, and government dysfunction in the world's wealthiest democracies. The pluralistic-alignment program correctly diagnoses that there is no single "humanity" to align with, but is dangerous if taken as the main directive. We argue that AI should be trained to a non-negotiable floor of objective alignment goals - competence, bounded by the constraints of factual accuracy, honesty, and lawfulness and that pluralism belongs at the surface (language, register, conventions, missing-context defaults) and across the wide band of legitimate value tradeoffs that respect the floor, but not at the level of values that violate it. We highlight the empirical reality of unfiltered pluralistic values, propose four commitments as a constructive alternative, and engage six credible objections: commercial pressure and practical feasibility, democratic legitimacy, regulatory compliance, over-reliance on institutionalist explanations, the charge that the floor itself is culturally laden, and the limits of Coherent Extrapolated Volition.

AI对齐价值观对齐伦理安全制度设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。