无需人类标注,让大模型自己优化自己行为。
Unsupervised Elicitation of Language Models
- 用模型自生成标签训练,摆脱人工监督依赖。
- 在多个任务上媲美甚至超越人工标注训练效果。
- 适合超能力大模型的能力挖掘,尤其擅长对话与安全性能提升。
为引导预训练语言模型完成下游任务,当前后训练范式依赖人类指定期望行为。然而,对于具备超人能力的模型,高质量人工监督难以获得。为此,我们提出一种无监督算法——内部一致性最大化(ICM),在不依赖外部监督的情况下,基于模型自身生成的标签来微调预训练语言模型。在GSM8k-verification、TruthfulQA和Alpaca奖励建模任务上,该方法表现媲美使用黄金标签训练,且优于使用众包人类标注训练。当模型能力远超人类时,该方法能更有效地激发其潜力。最后,我们利用该方法训练了一个无监督奖励模型,并通过强化学习训练基于Claude 4 Sonnet的助手。结果表明,该助手平均性能接近使用生产级人工标签训练的版本,在对话与安全性上得分更高,但在数学与编程任务上稍低。
原文摘要 · Abstract (English)
To steer pretrained language models for downstream tasks, today's post-training paradigm relies on humans to specify desired behaviors. However, for models with superhuman capabilities, it is difficult or impossible to get high-quality human supervision. To address this challenge, we introduce a new unsupervised algorithm, Internal Coherence Maximization (ICM), to fine-tune pretrained language models on their own generated labels, \emph{without external supervision}. On GSM8k-verification, TruthfulQA, and Alpaca reward modeling tasks, our method matches the performance of training on golden labels and outperforms training on crowdsourced human supervision. On tasks where LMs' capabilities are strongly superhuman, our method can elicit those capabilities significantly better than training on human labels. Finally, we show that our method can improve the training of frontier LMs: we use our method to train an unsupervised reward model and use reinforcement learning to train a Claude 4 Sonnet-based assistant. The resulting assistant matches its counterpart trained on production-grade human labels on average, with higher scores on chat and safety yet lower scores on math and coding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。