arXiv:2505.16927cs.CLcs.AI2025-05NeurIPS

让大模型自动发现并学习改进自身的隐含原则。

Latent Principle Discovery for Language Model Self-Improvement

  • 通过自校正机制挖掘模型内在的改进原则。
  • 7-8B模型迭代优化后,评分提升8-10%至23%。
  • 生成的规则可解释且多样,适合模型持续自我进化。

当用户希望提升语言模型生成质量时,明确具体的行为属性至关重要。然而,在多个领域中人工标注这些原则成本高昂。为此,我们提出一种方法,通过在自校正设置中显式建模,从模型自身挖掘引导其推理向人类偏好响应靠拢的潜在属性。该方法利用后验正则化的蒙特卡洛期望最大化算法,识别出最有效的少量隐含原则,并教会模型有策略地调用它们以内在优化输出。实验表明,多轮迭代下,7-8B参数的小型模型实现自提升:AlpacaEval胜率提升8-10%,MT-Bench平均得分提高0.3,IFEval上原则遵循胜率增长19-23%。聚类分析使发现的原则具备可解释性和多样性,同时保持性能。结果证明了自动化、原则驱动的后训练方案在持续自我改进中的潜力。

原文摘要 · Abstract (English)

When language model (LM) users aim to improve the quality of its generations, it is crucial to specify concrete behavioral attributes that the model should strive to reflect. However, curating such principles across many domains, even non-exhaustively, requires a labor-intensive annotation process. To automate this process, we propose eliciting these latent attributes that guide model reasoning toward human-preferred responses by explicitly modeling them in a self-correction setting. Our approach mines new principles from the LM itself and compresses the discovered elements to an interpretable set via clustering. Specifically, we employ a form of posterior-regularized Monte Carlo Expectation-Maximization to both identify a condensed set of the most effective latent principles and teach the LM to strategically invoke them in order to intrinsically refine its responses. We demonstrate that bootstrapping our algorithm over multiple iterations enables smaller language models (7-8B parameters) to self-improve, achieving +8-10% in AlpacaEval win-rate, an average of +0.3 on MT-Bench, and +19-23% in principle-following win-rate on IFEval. We also show that clustering the principles yields interpretable and diverse model-generated constitutions while retaining model performance. The gains that our method achieves highlight the potential of automated, principle-driven post-training recipes toward continual self-improvement.

自提升原则挖掘语言模型自校正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。