AI理解自我与他人需构建多元世界模型,而非强行统一认知。
Toward AI That Understands Self and Others: A World-Model Theory of Cognitive Diversity and Alignment
- 提出多阶段推断机制,让AI在有限条件下构建可预测的近似足够统计量。
- 揭示认知分歧源于信息处理差异,而非单纯价值观对立。
- 适合研究认知多样性、智能对齐与跨主体AI协作的学者。
现代社会信息爆炸却未形成统一认知,相同事件可被解读为自由、危险、不公等不同意义。现有讨论常将分歧归因于价值观或信念冲突,本文认为分歧已是后期现象。核心观点是:观察不等于推断——并非所有观察都具推断相关性,也非所有可观测对象都能成为估计目标。只有当能构建出对预测、评估或行动近似充分的状态表示时,该目标才可被接纳。论文基于此提出世界模型理论,重构认知为在信息、表征、观测和行动约束下构建近似充分统计量的过程,提出多阶段推断假设(MIA)与机制(MIM)。引入对齐映射与变换损失,分析异质世界模型如何在不坍缩为单一表征的前提下进行通信。因此,世界模型对齐的本质是可处理性:设计能使多元智能保持相互可处理且保留各自误差检测能力的AI系统。
原文摘要 · Abstract (English)
Modern societies possess more information than ever before, yet they do not converge toward a single shared understanding. The same events, facts, laws, technologies, or risks can be interpreted as evidence of freedom, danger, exclusion, injustice, responsibility, or unrealized possibility. Existing discussions often treat such disagreement as a conflict of values, preferences, or beliefs. This paper argues that disagreement is already a late-stage phenomenon. The central premise is simple but not trivial: observation is not yet inference. Not every observation becomes inferentially relevant, and not every possible object in an observation sequence becomes an estimation target. A possible target becomes admissible only when a state representation can be constructed that is approximately sufficient for prediction, evaluation, or action with respect to that target. This paper develops a world-model theory of cognitive diversity and alignment by reconstructing recognition as the construction of such approximate sufficient statistics under finite informational, representational, observational, and action constraints. It formulates this position as the Multi-Phase Inference Assumption (MIA) and defines its core internal mechanism as the Multi-Phase Inference Mechanism (MIM). The framework introduces alignment maps and transformation loss to analyze how heterogeneous world models communicate without being collapsed into a single representation. World-model alignment is therefore processability, not agreement: the design of AI systems that help heterogeneous forms of intelligence remain mutually processable while preserving their distinct error-detection capacities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。