大模型自进化:用自身生成数据并持续优化,摆脱对人工反馈的依赖。
Self-Improvement of Large Language Models: A Technical Overview and Future Outlook
- 构建闭环系统,让模型自主完成数据生成、选择、优化和推理改进。
- 通过自评估机制持续监控进展,驱动多阶段迭代提升性能。
- 适合关注AI自我演化、自动化模型升级的研究者与开发者。
随着大语言模型(LLMs)持续发展,仅依靠人工监督进行改进的成本越来越高且难以扩展。当模型在某些领域接近人类水平时,人工反馈提供的信息量已不足以支撑进一步优化。与此同时,模型自主决策和执行复杂任务的能力不断增强,使得模型开发过程中的多个环节可逐步实现自动化。这催生了自进化研究热潮——模型能够自主生成数据、评估输出,并迭代优化自身能力。本文从系统层面出发,提出统一框架,将自进化过程视为一个由四个紧密耦合阶段构成的闭环生命周期:数据获取、数据筛选、模型优化与推理精炼,并引入自主评估层。在此框架下,模型在各阶段均发挥核心作用:生成或收集数据、筛选有效信号、更新参数、优化输出,而自评估层则持续监测进展并指导跨阶段改进。基于该生命周期视角,我们系统梳理并分析了各组件的代表性技术方法,探讨当前局限性,并展望未来实现完全自进化大模型的研究方向。
原文摘要 · Abstract (English)
As large language models (LLMs) continue to advance, improving them solely through human supervision is becoming increasingly costly and limited in scalability. As models approach human-level capabilities in certain domains, human feedback may no longer provide sufficiently informative signals for further improvement. At the same time, the growing ability of models to make autonomous decisions and execute complex actions naturally enables abstractions in which components of the model development process can be progressively automated. Together, these challenges and opportunities have driven increasing interest in self-improvement, where models autonomously generate data, evaluate outputs, and iteratively refine their own capabilities. In this paper, we present a system-level perspective on self-improving language models and introduce a unified framework that organizes existing techniques. We conceptualize the self-improvement system as a closed-loop lifecycle, consisting of four tightly coupled processes: data acquisition, data selection, model optimization, and inference refinement, along with an autonomous evaluation layer. Within this framework, the model itself plays a central role in driving each stage: collecting or generating data, selecting informative signals, updating its parameters, and refining outputs, while the autonomous evaluation layer continuously monitors progress and guides the improvement cycle across stages. Following this lifecycle perspective, we systematically review and analyze representative methods for each component from a technical standpoint. We further discuss current limitations and outline our vision for future research toward fully self-improving LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。