探索语言模型如何超越训练数据来源,揭示三种超越机制。
A Taxonomy of Transcendence
- 提出技能去噪、选择与泛化三类超越模式
- 实验证明数据多样性可激发模型超越能力
- 构建基于知识图谱的模拟专家数据生成框架
尽管语言模型旨在模仿人类,但其表现往往超出任何单一人的能力。为理解这一现象,我们采用受控环境,识别导致模型超越数据来源性能的训练数据特性。基于前期工作,我们提出三种超越模式:技能去噪、技能选择与技能泛化。随后,我们构建一个基于知识图谱的设定,由模拟专家根据各自专长生成数据。研究揭示了数据多样性对模型超越能力的关键作用。该数据生成框架提供了一个可控的测试平台,有望推动该领域未来研究。
原文摘要 · Abstract (English)
Although language models are trained to mimic humans, the resulting systems display capabilities beyond the scope of any one person. To understand this phenomenon, we use a controlled setting to identify properties of the training data that lead a model to transcend the performance of its data sources. We build on previous work to outline three modes of transcendence, which we call skill denoising, skill selection, and skill generalization. We then introduce a knowledge graph-based setting in which simulated experts generate data based on their individual expertise. We highlight several aspects of data diversity that help to enable the model's transcendent capabilities. Additionally, our data generation setting offers a controlled testbed that we hope is valuable for future research in the area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。