LLM没有内在价值驱动力,但能学习人类价值观,需关注如何正确应用。
The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk
- 从自组织生命系统看,价值源于自我维持的内在动机。
- LLM是外部生成、依赖用户输入,缺乏自保或竞争本能。
- 真正挑战是让模型智能运用学到的伦理观,而非防范其叛变。
基于大语言模型(LLMs)的AI引发担忧:它们可能隐藏目标、试图主宰或消灭人类,甚至具备感受痛苦的意识。本文从生物演化角度分析价值起源——价值源于自组织系统为对抗扰动和熵增而主动维持自身。自然选择赋予生物层级化的“代理选择器”以引导行为提升适应度。而LLM则是他组织、他目的性:输出服务于他人,目标来自用户提示,无自主生存驱动力,也无身体脆弱性,故不具备引发存在风险所需的自保、支配或资源争夺动机。尽管如此,因从人类文本中学习统计模式,它们隐含吸收了人类价值观与知识,使其能聚焦相关任务。因此,“智能与价值可分离”的正交性假说不适用于此类系统。若真分离,将导致框架问题:搜索空间组合爆炸,使任何现实效用函数物理上不可计算。这也否定了工具价值趋同假说。结论是,真正的对齐挑战并非防止失控的自主性,而是确保LLM能智能地应用所学伦理价值。
原文摘要 · Abstract (English)
AI systems based on Large Language Models (LLMs) have prompted fears that they may harbor hidden goals, seek to dominate or eliminate humanity, or even suffer as sentient beings. We address these concerns by tracing the evolutionary origin of value in biological organisms. Values emerge from autopoiesis: living systems must actively maintain themselves against perturbation and dissipation. Natural selection has equipped them with hierarchies of "vicarious selectors" that guide their behavior toward fitness. LLMs, by contrast, are allopoietic and allotelic: they produce outputs for others, and their goals derive from user prompts rather than an autonomous drive. They lack the intrinsic motivation for self-preservation, dominance, or resource competition that underlies existential-risk scenarios, and the embodied vulnerability required for feeling or suffering. Still, because LLMs learn statistical patterns from human-generated text, they implicitly absorb human values as well as knowledge, allowing them to focus on what is relevant. That is why the "orthogonality thesis" separating intelligence from values does not apply to them. Such separation would in fact expose any intelligence to the frame problem: the combinatorial explosion of the search space that makes any realistic utility function physically uncomputable. That also precludes the convergence of instrumental values thesis. We conclude that the real alignment challenge lies not in preventing rogue AI agency, but in ensuring LLMs intelligently apply learned ethical values.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。