用信息论方法让大模型同时满足多种人类偏好,提升可控性与对齐度。
Multi-Objective Exploration and Preference Optimization via Mutual Information
- 通过最大化响应、反馈与偏好向量间的互信息,统一探索与对齐过程
- 在安全对齐与助手指令任务中,响应与偏好向量对齐度显著提升
- 适合需要多目标平衡的对话系统、价值观对齐等场景
将大型语言模型与多样且异构的人类价值观对齐,需采用多目标对齐方法以有效权衡冲突的偏好维度。现有方法通过条件化偏好向量的策略训练,并利用在线直接偏好优化实现权衡。然而,探索不确定性会导致不同偏好向量生成的响应奖励分布重叠,使生成结果无法有效对齐对应偏好向量。本文提出基于互信息的多目标探索与偏好优化(MI-EPO),一种信息论框架。该框架通过最大化生成响应、偏好反馈与偏好向量之间的联合条件互信息,统一多目标探索与对齐。结合概率路由机制,MI-EPO自然分解目标对齐与偏好感知探索,促使模型生成可区分且与不同偏好条件对齐的响应。在安全对齐与助手指令任务上的实验表明,MI-EPO显著提升生成响应与偏好向量的对齐度,增强输出可控性,并在多个目标间实现稳定权衡。
原文摘要 · Abstract (English)
Aligning large language models with diverse and heterogeneous human values requires multi-objective alignment methods to effectively trade off conflicting preference dimensions. Current methods achieve this trade-off by training policies conditioned on preference vectors and leveraging online direct preference optimization. However, exploration uncertainty can cause the reward distributions of responses generated under different preference vectors to overlap, and the generated responses may fail to effectively align with the corresponding preference vectors. In this paper, we propose Multi-Objective Exploration and Preference Optimization via Mutual Information (MI-EPO), an information-theoretic framework. It unifies multi-objective exploration and alignment by maximizing the joint conditional mutual information among generated responses, preference feedback, and preference vectors. By incorporating a probabilistic routing mechanism, MI-EPO naturally decomposes objective alignment and preference-aware exploration, encouraging the model to generate responses that are distinguishable and aligned with different preference conditions. Experiments on safe alignment and helpful assistant tasks show that MI-EPO significantly improves the alignment between generated responses and preference vectors, makes the outputs more controllable, and achieves stable trade-offs across multiple objectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。