将模型训练改为价值塑造,从一开始就培养模型的内在价值观。
From Model Training to Model Raising
- 用第一人称视角重构训练数据,让模型以体验方式学习
- 通过模拟社交互动和分阶段数据排序,早期建立价值认同
- 适合关注模型伦理对齐与长期安全的研究者
当前的AI训练方法在模型核心能力形成后才进行价值观对齐,导致模型易产生偏差且缺乏深层价值体系。本文提出从‘模型训练’转向‘模型养育’的新范式,将对齐融入模型发展的全过程。核心在于重新设计训练语料:将数据重构为第一人称视角、将信息转化为生活经验、模拟社会互动,并分层组织训练顺序。我们预期这一重构将使模型从第一个训练标记起就内嵌价值承诺,使得知识、技能与价值观难以分离。在大语言模型在诸多任务上已超越人类能力的背景下,这种早期价值锚定显得尤为关键。
原文摘要 · Abstract (English)
Current AI training methods align models with human values only after their core capabilities have been established, resulting in models that are easily misaligned and lack deep-rooted value systems. We propose a paradigm shift from "model training" to "model raising", in which alignment is woven into a model's development from the start. We identify several key components for this paradigm, all centered around redesigning the training corpus: reframing training data from a first-person perspective, recontextualizing information as lived experience, simulating social interactions, and scaffolding the ordering of training data. We expect that this redesign of the training corpus will lead to an early commitment to values from the first training token onward, such that knowledge, skills, and values are intrinsically much harder to separate. In an ecosystem in which large language model capabilities start overtaking human capabilities in many tasks, this seems to us like a critical need.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。