arXiv:2412.11145cs.CL2024-12被引 2

探索强于人类的模型如何对齐人类价值观,提出超对齐新框架。

The Superalignment of Superhuman Intelligence with Large Language Models

  • 提出超对齐三模块框架:攻击者、学习者、评论者协同进化。
  • 在人类无法标注复杂任务时,实现从噪声标签中高效学习。
  • 适合研究通用人工智能安全与对齐的学者参考。

得益于大语言模型和多模态语言模型的快速发展,我们已见证超人类智能的出现。随着这类超人类模型应用日益广泛,一个关键问题浮现:如何确保这些模型仍安全、可靠且与人类价值观对齐?本文从学习视角探讨超对齐概念,梳理了从大规模预训练、监督微调到对齐训练的学习范式转变。定义超对齐为:当任务复杂到人类专家难以标注、模型能力超越人类时,设计可扩展的高效对齐算法,从带噪声的标签数据(点级样本或成对偏好数据)中学习。本文指出超对齐的关键研究问题:弱到强泛化、可扩展监督与评估。提出一个概念框架,包含三个模块:攻击者生成对抗性查询以暴露学习者弱点;学习者通过批判模型生成的可扩展反馈及少量人类专家指导自我优化;批判者针对查询-响应对生成批评或解释,以改进学习者。讨论各模块中的重要研究问题,并关联自对齐、自博弈、自精炼等前沿思路。最后展望未来方向,包括新兴风险识别与多维度对齐。

原文摘要 · Abstract (English)

We have witnessed superhuman intelligence thanks to the fast development of large language models and multimodal language models. As the application of such superhuman models becomes more and more popular, a critical question arises here: how can we ensure superhuman models are still safe, reliable and aligned well to human values? In this position paper, we discuss the concept of superalignment from the learning perspective to answer this question by outlining the learning paradigm shift from large-scale pretraining, supervised fine-tuning, to alignment training. We define superalignment as designing effective and efficient alignment algorithms to learn from noisy-labeled data (point-wise samples or pair-wise preference data) in a scalable way when the task becomes very complex for human experts to annotate and the model is stronger than human experts. We highlight some key research problems in superalignment, namely, weak-to-strong generalization, scalable oversight, and evaluation. We then present a conceptual framework for superalignment, which consists of three modules: an attacker which generates adversary queries trying to expose the weaknesses of a learner model; a learner which will refine itself by learning from scalable feedbacks generated by a critic model along with minimal human experts; and a critic which generates critics or explanations for a given query-response pair, with a target of improving the learner by criticizing. We discuss some important research problems in each component of this framework and highlight some interesting research ideas that are closely related to our proposed framework, for instance, self-alignment, self-play, self-refinement, and more. Last, we highlight some future research directions for superalignment, including identification of new emergent risks and multi-dimensional alignment.

对齐大模型安全自迭代

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。