arXiv:2512.05464cs.CLcs.AI2025-12被引 1

提出动态对齐框架,让大模型自我改进并达成更全面的智能体协作目标。

Dynamic Alignment for Collective Agency: Toward a Scalable Self-Improving Framework for Open-Ended LLM Alignment

  • 用自生成数据与自奖励机制实现大模型自我对齐
  • 在保持通用NLP能力前提下成功对齐集体代理目标
  • 适合关注大模型自主进化与高级对齐的研究者

大型语言模型(LLMs)通常通过人类偏好数据或预设原则(如有用性、诚实性、无害性)进行对齐。然而,随着人工智能向通用智能(AGI)和超智能(ASI)发展,这些价值体系可能不再足够。此外,基于人类反馈的对齐方法资源消耗大且难以扩展。尽管已有研究探索以AI反馈为基础的自改进对齐方法作为可扩展替代方案,但其大多局限于传统对齐价值。本文探索更全面的对齐目标与可扩展的自改进对齐路径。为超越传统对齐范式,我们引入集体代理(Collective Agency, CA)——一种统一且开放的对齐价值,旨在促进集成的智能体能力。我们提出动态对齐(Dynamic Alignment)框架,使大模型能够迭代式自我对齐。该框架包含两个核心组件:(1)由大模型自动生成训练数据;(2)自奖励机制,即策略模型评估自身输出候选并赋予奖励,用于基于GRPO的持续学习。实验表明,该方法在保持通用NLP能力的同时,成功将模型对齐至集体代理目标。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are typically aligned with human values using preference data or predefined principles such as helpfulness, honesty, and harmlessness. However, as AI systems progress toward Artificial General Intelligence (AGI) and Artificial Superintelligence (ASI), such value systems may become insufficient. In addition, human feedback-based alignment remains resource-intensive and difficult to scale. While AI-feedback-based self-improving alignment methods have been explored as a scalable alternative, they have largely remained constrained to conventional alignment values. In this work, we explore both a more holistic alignment objective and a scalable, self-improving alignment approach. Aiming to transcend conventional alignment norms, we introduce Collective Agency (CA)-a unified and open-ended alignment value that encourages integrated agentic capabilities. We also propose Dynamic Alignment-an alignment framework that enables an LLM to iteratively align itself. Dynamic Alignment comprises two key components: (1) automated training dataset generation with LLMs, and (2) a self-rewarding mechanism, where the policy model evaluates its own output candidates and assigns rewards for GRPO-based learning. Experimental results demonstrate that our approach successfully aligns the model to CA while preserving general NLP capabilities.

大模型对齐自改进集体代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。