arXiv:2601.01665cs.LGcs.AI2026-01

提出对抗生成与防御训练框架,提升多目标组合优化模型的鲁棒性。

Adversarial Instance Generation and Robust Training for Neural Combinatorial Optimization with Multiple Objectives

  • 基于偏好设计对抗攻击,生成暴露模型弱点的难例
  • 在MOTSP/MOCVRP/MOKP上验证,攻防策略显著提升泛化能力
  • 适合关注神经优化器稳定性与抗干扰能力的研究者

深度强化学习(DRL)在解决多目标组合优化问题(MOCOPs)方面展现出巨大潜力,但其鲁棒性仍不足,尤其在多样复杂的分布下。本文提出统一的面向鲁棒性的偏好条件化DRL求解框架。设计基于偏好的对抗攻击方法,生成能暴露求解器弱点的难例,并通过帕累托前沿质量下降程度量化攻击影响。进一步引入将难度感知偏好选择融入对抗训练的防御策略,降低对有限偏好区域的过拟合,提升分布外性能。在多目标旅行商问题(MOTSP)、多目标容量车辆路径问题(MOCVRP)和多目标背包问题(MOKP)上的实验表明,所提攻击方法能有效生成针对不同求解器的难例;所提防御方法显著增强神经求解器的鲁棒性与泛化性,在困难或分布外实例上表现更优。

原文摘要 · Abstract (English)

Deep reinforcement learning (DRL) has shown great promise in addressing multi-objective combinatorial optimization problems (MOCOPs). Nevertheless, the robustness of these learning-based solvers has remained insufficiently explored, especially across diverse and complex problem distributions. In this paper, we propose a unified robustness-oriented framework for preference-conditioned DRL solvers for MOCOPs. Within this framework, we develop a preference-based adversarial attack to generate hard instances that expose solver weaknesses, and quantify the attack impact by the resulting degradation on Pareto-front quality. We further introduce a defense strategy that integrates hardness-aware preference selection into adversarial training to reduce overfitting to restricted preference regions and improve out-of-distribution performance. The experimental results on multi-objective traveling salesman problem (MOTSP), multi-objective capacitated vehicle routing problem (MOCVRP), and multi-objective knapsack problem (MOKP) verify that our attack method successfully learns hard instances for different solvers. Furthermore, our defense method significantly strengthens the robustness and generalizability of neural solvers, delivering superior performance on hard or out-of-distribution instances.

强化学习组合优化对抗训练多目标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。