arXiv:2508.09016cs.CLcs.LG2025-08EMNLP综述被引 4

不微调就能让大模型更符合人类价值观,这篇综述系统梳理了三种无训练对齐方法。

A Survey on Training-free Alignment of Large Language Models

  • 通过提示工程、解码时调整和生成后修正实现无训练对齐
  • 覆盖预解码、解码中、解码后三个阶段的主流技术方案
  • 适合资源受限或无法访问模型的场景,助力安全可控的AI应用

大型语言模型(LLM)的对齐旨在确保其输出符合人类价值观、伦理标准和法律规范。传统对齐方法多依赖计算资源密集的微调(FT),可能引发知识退化,且在模型不可访问或算力受限时难以应用。相比之下,无训练(TF)对齐技术——利用上下文学习、解码时调整和生成后修正——无需大量重训练即可实现对齐,适用于开源与闭源环境。本文首次系统性综述了TF对齐方法,按预解码、解码中、解码后三阶段分类,并从LLM与多模态LLM(MLLM)视角深入分析其机制与局限。同时识别关键挑战与未来方向,推动更普惠、高效的无训练对齐技术发展。通过整合快速增长的研究成果,本综述为实践者提供指引,促进更安全可靠的大型语言模型演进。

原文摘要 · Abstract (English)

The alignment of large language models (LLMs) aims to ensure their outputs adhere to human values, ethical standards, and legal norms. Traditional alignment methods often rely on resource-intensive fine-tuning (FT), which may suffer from knowledge degradation and face challenges in scenarios where the model accessibility or computational resources are constrained. In contrast, training-free (TF) alignment techniques--leveraging in-context learning, decoding-time adjustments, and post-generation corrections--offer a promising alternative by enabling alignment without heavily retraining LLMs, making them adaptable to both open-source and closed-source environments. This paper presents the first systematic review of TF alignment methods, categorizing them by stages of pre-decoding, in-decoding, and post-decoding. For each stage, we provide a detailed examination from the viewpoint of LLMs and multimodal LLMs (MLLMs), highlighting their mechanisms and limitations. Furthermore, we identify key challenges and future directions, paving the way for more inclusive and effective TF alignment techniques. By synthesizing and organizing the rapidly growing body of research, this survey offers a guidance for practitioners and advances the development of safer and more reliable LLMs.

大模型对齐无训练LLM安全综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。