arXiv:2608.29978cs.CLcs.LG2026-08中稿 · EMNLP

用进化算法训练专家混合模型,实现大模型生成的精细调控。

Evolutionary Soups: Evolving Mixture-of-Experts for Multi-Objective LLM Alignment

论文配图:Evolutionary Soups: Evolving Mixture-of-Experts for Multi-Objective LLM Alignment
图 1 · 摘自论文原文
  • 通过进化算法优化门控网络,动态生成专家融合系数。
  • 在三类任务中均提升约20%的多目标评估指标表现。
  • 适合需要灵活控制生成风格的大模型应用场景。

大型语言模型日益需要生成满足多个竞争目标的响应。由于最优权衡依赖于用户偏好和输入提示,可控的多目标生成必须在推理时动态调整模型而无需重新训练。为此,我们提出进化汤(Evolutionary Soups),一种基于专家混合(MoE)框架的细粒度生成控制方法,其门控网络通过进化算法训练。每层门控网络从隐藏状态表示中动态生成专家融合系数,进化算法引入贪心超体积贡献机制,有效演化门控网络,在大规模且嘈杂的数据集上持续提升性能,并更全面覆盖非凸帕累托前沿。在三个任务上的实验表明,该方法在所有任务中均优于基线:在超体积、线性效用和切比雪夫效用上取得最佳表现,相比可控方法提升约20%。

原文摘要 · Abstract (English)

Large language models are increasingly required to generate responses that satisfy multiple competing objectives. Since optimal trade-offs depend on both user preferences and input prompts, controllable multi-objective generation must dynamically adapt models at inference time without retraining. To address this, we propose Evolutionary Soups, a mixture-of-experts framework for fine-grained generation control, with gating networks trained via an evolutionary algorithm. The per-layer gating networks dynamically produce expert-merging coefficients from hidden-state representations, while the evolutionary algorithm incorporates greedy hypervolume contribution for effective evolution of these gating networks, achieving consistent improvements on large and noisy training datasets and broader coverage of the non-convex Pareto front. Experiments across three tasks demonstrate the effectiveness of Evolutionary Soups over baselines: it achieves the best hypervolume, linear utility, and Tchebyshev utility (~20% improvement) among controllable methods on all tasks.

专家混合多目标优化进化算法生成控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。