arXiv:2505.20925cs.CLcs.AI2025-05被引 15

无需训练即可让大模型同时满足多种用户偏好,灵活适配不同需求。

Multi-objective Large Language Model Alignment with Hierarchical Experts

  • 采用分层专家架构,通过轻量级组件实现参数高效对齐。
  • 在6个基准上测试14项任务、200种偏好,性能超越15个基线方法。
  • 适合需要快速适配多样需求的场景,如个性化对话系统。

大语言模型同时满足多个目标的对齐仍是重大挑战,尤其因人类偏好多样且常存在冲突。现有方法难以有效平衡权衡,往往需代价高昂的重新训练或在偏好帕累托前沿上表现不佳。本文提出HoE(分层专家混合模型),一种轻量、参数高效且即插即用的方法,无需模型训练即可使大模型适应整个帕累托前沿,满足多样化用户偏好。HoE包含三层结构:LoRA专家、路由器专家与偏好路由机制,在参数量、训练成本与性能间取得最优权衡。我们在6个基准上对14项任务和200种不同偏好进行了评估,结果表明其性能显著优于15个近期基线方法。代码已提供于补充材料中。

原文摘要 · Abstract (English)

Aligning large language models (LLMs) to simultaneously satisfy multiple objectives remains a significant challenge, especially given the diverse and often conflicting nature of human preferences. Existing alignment methods struggle to balance trade-offs effectively, often requiring costly retraining or yielding suboptimal results across the Pareto frontier of preferences. In this paper, we introduce \textit{HoE}(Hierarchical Mixture-of-Experts), a \textit{lightweight}, \textit{parameter-efficient}, and \textit{plug-and-play} approach that eliminates the need for model training, while enabling LLMs to adapt across the entire Pareto frontier and accommodate diverse user preferences. In particular, \textit{HoE} consists of three hierarchical components: LoRA Experts, Router Experts and Preference Routing, reaching optimal Pareto frontiers and achieving a trade-off between parameter size, training cost, and performance. We evaluate \textit{HoE} across various tasks on 14 objectives and 200 different preferences among 6 benchmarks, demonstrating superior performance over 15 recent baselines. Code is available in the supplementary materials.

大模型对齐多目标优化专家混合参数效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。