arXiv:2502.05475cs.LG2025-02被引 14

数据结构如何塑造模型行为,是AI对齐的关键。

You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation

  • 同一训练集上表现相似的模型,内部机制可能完全不同。
  • 模型泛化能力差异源于数据与模型结构的深层关联。
  • 适合关注AI安全与可解释性的研究者阅读。

本文主张理解数据分布中的结构与训练后模型内部结构之间的关系,是实现人工智能对齐的核心。首先,我们指出两个神经网络在训练集上表现相当,却可能以本质不同的方式计算输出,从而导致不同的泛化行为。因此,仅依赖标准测试与评估不足以确保广泛部署的通用智能系统安全性。为推动从评估迈向坚实的数学科学,我们需要建立统计基础,深入理解数据分布结构、模型内部结构及其如何共同决定泛化性能。

原文摘要 · Abstract (English)

In this position paper, we argue that understanding the relation between structure in the data distribution and structure in trained models is central to AI alignment. First, we discuss how two neural networks can have equivalent performance on the training set but compute their outputs in essentially different ways and thus generalise differently. For this reason, standard testing and evaluation are insufficient for obtaining assurances of safety for widely deployed generally intelligent systems. We argue that to progress beyond evaluation to a robust mathematical science of AI alignment, we need to develop statistical foundations for an understanding of the relation between structure in the data distribution, internal structure in models, and how these structures underlie generalisation.

AI对齐模型泛化数据结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。