arXiv:2507.08967cs.CL2025-07被引 4

无需外部标注,模型自动生成对比样本实现动态对齐

Self-Improving Model Steering

  • 通过迭代自改进生成并优化对比样本
  • 在多个基准上显著优于现有方法
  • 适合需要灵活适配场景的LLM对齐应用

模型对齐是一种在推理时动态调整大语言模型以匹配人类偏好的强大技术。然而,传统方法高度依赖外部标注数据,不仅限制了其在不同上下文中的适应能力,还使其效果受标注质量制约。本文提出SIMS,首个无需外部监督的自改进模型对齐框架。其核心是通过迭代自改进循环,自主生成并优化对比样本,实现自适应、上下文相关的对齐。此外,SIMS采用提示排序和对比采样等新策略,进一步提升对齐效果。在多种大模型和基准上的广泛评估表明,SIMS在对齐效果和适应性方面均显著优于现有方法,凸显了自改进模型对齐作为未来推理时大模型对齐研究的前景。

原文摘要 · Abstract (English)

Model steering represents a powerful technique that dynamically aligns large language models (LLMs) with human preferences during inference. However, conventional model-steering methods rely heavily on externally annotated data, not only limiting their adaptability to varying contexts but also tethering their effectiveness to annotation quality. In this paper, we present SIMS, the first self-improving model-steering framework that operates without relying on external supervision. At its core, SIMS autonomously generates and refines contrastive samples through iterative self-improvement cycles, enabling adaptive, context-specific steering. Additionally, SIMS employs novel strategies, including prompt ranking and contrast sampling, to further enhance steering efficacy. Extensive evaluation across diverse LLMs and benchmarks demonstrates that SIMS substantially outperforms existing methods in steering effectiveness and adaptability, highlighting self-improving model steering as a promising direction for future research on inference-time LLM alignment.

模型对齐自改进LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。