用多轮反馈让大模型协同优化启发式算法,提升组合优化性能
ReVEL: Multi-Turn Reflective LLM-Guided Heuristic Evolution via Structured Performance Feedback

- 将启发式算法分组,通过多轮大模型反思式迭代优化
- 在多个基准测试中显著优于现有方法,性能提升稳定
- 适合需要自动设计优化算法的研究者和工程师
为解决NP难组合优化问题设计高效启发式算法仍具挑战,通常需大量领域知识。近期基于大模型的进化方法在自动化启发式生成方面展现出潜力,但多数方法仅独立优化或依赖有限的成对反馈。我们提出ReVEL:基于结构化性能反馈的多轮反思式大模型引导启发式演化框架,支持群体化、多轮的启发式协同优化。ReVEL将启发式算法按行为特征划分为感知型反思组,包括基于相似性的局部优化组与基于多样性的探索搜索组。每组内大模型利用累积性能反馈进行多轮迭代优化,实现相关启发式算法的联合分析与渐进改进。在标准组合优化基准上的实验表明,ReVEL在多种设置和大模型基底上均显著优于现有大模型引导进化基线。额外分析显示,行为感知分组有助于在迭代演化过程中保持更一致的优化轨迹。
原文摘要 · Abstract (English)
Designing effective heuristics for NP-hard combinatorial optimization problems remains challenging and often requires substantial domain expertise. Recent LLM-guided evolutionary methods have shown promise for automated heuristic generation, but most existing approaches refine heuristics independently or through limited pairwise feedback. We propose ReVEL: Multi-Turn Reflective LLM-Guided Heuristic Evolution via Structured Performance Feedback, a framework for group-wise multi-turn heuristic refinement. ReVEL organizes heuristics into behavior-aware reflective groups, including similarity-driven groups for localized refinement and diversity-driven groups for exploratory search. Within each group, the LLM performs iterative multi-turn refinement using accumulated performance feedback, enabling related heuristics to be jointly analyzed and progressively improved across evolutionary iterations. Experiments on standard combinatorial optimization benchmarks show that ReVEL generally improves optimization performance over existing LLM-guided evolutionary baselines across multiple settings and LLM backbones. Additional analyses suggest that behavior-aware grouping contributes to more consistent refinement trajectories during iterative heuristic evolution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。