让大模型自我进化,同时懂用户偏好和专业领域知识。
A Novel Self-Evolution Framework for Large Language Models
- 分两阶段优化:先学领域知识,再调用户偏好。
- 在多个任务上超越传统微调和记忆增强方法。
- 适合需要持续更新能力的智能对话系统。
大语言模型的能力受限于预训练,虽可通过后训练优化,但现有策略如基于记忆的检索或偏好优化,难以提升模型的领域认知。为此,我们提出双阶段自演化(DPSE)框架,同步优化用户偏好适应与领域专长。DPSE引入甄别模块,提取多维交互信号并估算满意度,指导主题感知与偏好驱动的数据扩展。扩增数据支持两阶段微调:监督式领域对齐,随后频率感知的偏好优化。在通用NLP基准与长期对话任务中,DPSE始终优于监督微调、偏好优化及记忆增强基线。消融实验证明各模块有效性。该框架为大模型的持续自主演化提供新路径。
原文摘要 · Abstract (English)
The capabilities of Large Language Models (LLMs) are limited to some extent by pre-training, so some researchers optimize LLMs through post-training. Existing post-training strategies, such as memory-based retrieval or preference optimization, improve user alignment yet fail to enhance the model's domain cognition. To bridge this gap, we propose a novel Dual-Phase Self-Evolution (DPSE) framework that jointly optimizes user preference adaptation and domain-specific competence. DPSE introduces a Censor module to extract multi-dimensional interaction signals and estimate satisfaction scores, which guide structured data expansion via topic-aware and preference-driven strategies. These expanded datasets support a two-stage fine-tuning pipeline: supervised domain grounding followed by frequency-aware preference optimization. Experiments across general NLP benchmarks and long-term dialogue tasks demonstrate that DPSE consistently outperforms Supervised Fine-Tuning, Preference Optimization, and Memory-Augmented baselines. Ablation studies validate the contribution of each module. In this way, our framework provides an autonomous path toward continual self-evolution of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。