arXiv:2506.08898cs.AI2025-06NeurIPS被引 7

根据偏好选择模型结构,提升多目标组合优化效果

Preference-Driven Multi-Objective Combinatorial Optimization with Conditional Computation

  • 按子问题动态分配专用神经架构,实现条件化计算
  • 基于胜败解对的偏好学习,无需显式奖励值
  • 在4个经典测试集上显著优于现有方法,通用性强

近期深度强化学习方法通过将多目标组合优化问题(MOCOPs)分解为多个与特定权重向量相关的子问题,取得了显著进展。然而,这些方法通常同等对待所有子问题,并用单一模型求解,限制了对解空间的有效探索,导致性能不佳。为此,我们提出POCCO——一种可即插即用的新框架,支持根据偏好信号自适应选择子问题的模型结构并进行优化。具体而言,设计了条件计算模块,将子问题路由至专用神经架构;同时提出偏好驱动优化算法,学习胜者与败者解之间的成对偏好。我们在两个先进的神经MOCOP方法上验证了POCCO的有效性与通用性,在四个经典MOCOP基准测试中均展现出显著优势和强泛化能力。

原文摘要 · Abstract (English)

Recent deep reinforcement learning methods have achieved remarkable success in solving multi-objective combinatorial optimization problems (MOCOPs) by decomposing them into multiple subproblems, each associated with a specific weight vector. However, these methods typically treat all subproblems equally and solve them using a single model, hindering the effective exploration of the solution space and thus leading to suboptimal performance. To overcome the limitation, we propose POCCO, a novel plug-and-play framework that enables adaptive selection of model structures for subproblems, which are subsequently optimized based on preference signals rather than explicit reward values. Specifically, we design a conditional computation block that routes subproblems to specialized neural architectures. Moreover, we propose a preference-driven optimization algorithm that learns pairwise preferences between winning and losing solutions. We evaluate the efficacy and versatility of POCCO by applying it to two state-of-the-art neural methods for MOCOPs. Experimental results across four classic MOCOP benchmarks demonstrate its significant superiority and strong generalization.

多目标优化强化学习条件计算偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。