让大模型学会按个人偏好调整行为,同时守住伦理底线。
A Survey on Personalized Alignment -- The Missing Piece for Large Language Models in Real-World Applications
- 提出个性化对齐统一框架,包含偏好记忆、个性生成和反馈修正
- 首次系统梳理个性化对齐方法,覆盖多种实际应用场景
- 适合关注大模型落地与伦理合规的研究者和开发者
大型语言模型(LLMs)展现出强大能力,但在真实应用中暴露出关键短板:无法在遵守普适人类价值观的前提下适应个体偏好。现有对齐技术采用‘一刀切’模式,难以满足用户多样背景与需求。本文首次全面综述个性化对齐——一种使LLM能在伦理边界内根据个体偏好调整行为的新范式。我们提出统一框架,涵盖偏好记忆管理、个性化生成与基于反馈的对齐,并系统分析实现路径,评估其在不同场景下的有效性。通过审视现有技术、潜在风险与未来挑战,本综述为构建更灵活且符合伦理的LLM提供结构化基础。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities, yet their transition to real-world applications reveals a critical limitation: the inability to adapt to individual preferences while maintaining alignment with universal human values. Current alignment techniques adopt a one-size-fits-all approach that fails to accommodate users' diverse backgrounds and needs. This paper presents the first comprehensive survey of personalized alignment-a paradigm that enables LLMs to adapt their behavior within ethical boundaries based on individual preferences. We propose a unified framework comprising preference memory management, personalized generation, and feedback-based alignment, systematically analyzing implementation approaches and evaluating their effectiveness across various scenarios. By examining current techniques, potential risks, and future challenges, this survey provides a structured foundation for developing more adaptable and ethically-aligned LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。