arXiv:2410.14184cs.CL2024-10NAACL被引 8

让大模型在推理时动态适应不同用户偏好,提升实用性。

MetaAlign: Align Large Language Models with Diverse Preferences during Inference Time

  • 提出在推理阶段动态对齐用户偏好的新方法MetaAlign。
  • 在自建数据集上训练的模型可即时响应各类显式/隐式偏好。
  • 适合需要个性化交互的应用场景,如客服、内容生成。

大型语言模型(LLMs)从海量文本中学习到丰富知识和强大能力,成为多种应用的强大工具。为提升可用性,使其与人类偏好对齐至关重要。现有对齐方法如基于人类反馈的强化学习(RLHF)和直接偏好优化(DPO)通常将预设偏好嵌入模型参数中,导致对齐结果静态化,难以应对实际应用中人类偏好的多样性。针对这一问题,本文提出一种新方法——MetaAlign,旨在使大模型在推理阶段能够动态对齐各种显式或隐式指定的偏好。实验表明,在精心构建的MetaAlign数据集上训练的模型,可在推理阶段有效响应任意指定偏好,验证了该方法的可行性。我们希望本工作能为语言模型对齐提供新思路。

原文摘要 · Abstract (English)

Large Language Models (LLMs) acquire extensive knowledge and remarkable abilities from extensive text corpora, making them powerful tools for various applications. To make LLMs more usable, aligning them with human preferences is essential. Existing alignment techniques, such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO), typically embed predefined preferences directly within the model's parameters. These methods, however, often result in a static alignment that can not account for the diversity of human preferences in practical applications. In response to this challenge, we propose an effective method, \textbf{MetaAlign}, which aims to help LLMs dynamically align with various explicit or implicit preferences specified at inference time. Experimental results show that LLMs optimized on our meticulously constructed MetaAlign Dataset can effectively align with any preferences specified at the inference stage, validating the feasibility of MetaAlign. We hope that our work can provide some insights into the alignment of language models.

大模型对齐推理时调整个性化生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。