arXiv:2505.24779cs.LG2025-05被引 2

提出新基准评估生成的混合整数规划实例是否真实有用。

Are Your Generated Instances Truly Useful? GenBench-MILP: A Benchmark Suite for MILP Instance Generation

  • 从数学有效性、结构相似性、计算难度等四维度评估生成实例质量。
  • 发现高结构相似度实例在求解器行为上差异巨大,计算难度不一。
  • 适合关注生成实例真实性和算法评估的研究者使用。

基于机器学习的混合整数线性规划(MILP)实例生成方法迅速发展,旨在提供多样化的训练数据集。然而,一个关键问题仍待解决:这些生成的实例是否真正有用且具有现实性?当前评估常依赖表面结构指标或简单的可解性检查,往往无法捕捉真实问题的计算复杂性。为此,我们提出GenBench-MILP,一个全面的基准套件,用于标准化、客观地评估MILP生成器。该框架从数学有效性、结构相似性、计算难度和下游任务实用性四个维度评估实例质量。其独特创新在于分析求解器内部特征——包括根节点间隙、启发式成功率和割平面使用率。将求解器的动态行为视为专家判断,揭示了静态图特征所忽略的细微计算差异。实验表明,即使结构相似度高的实例,其与求解器的交互方式和难度水平也可能截然不同。GenBench-MILP为严格比较和推动高保真实例生成器的发展提供了多维评估工具。

原文摘要 · Abstract (English)

The proliferation of machine learning-based methods for Mixed-Integer Linear Programming (MILP) instance generation has surged, driven by the need for diverse training datasets. However, a critical question remains: Are these generated instances truly useful and realistic? Current evaluation protocols often rely on superficial structural metrics or simple solvability checks, which frequently fail to capture the true computational complexity of real-world problems. To bridge this gap, we introduce GenBench-MILP, a comprehensive benchmark suite designed for the standardized and objective evaluation of MILP generators. Our framework assesses instance quality across four key dimensions: mathematical validity, structural similarity, computational hardness, and utility in downstream tasks. A distinctive innovation of GenBench-MILP is the analysis of solver-internal features -- including root node gaps, heuristic success rates, and cut plane usage. By treating the solver's dynamic behavior as an expert assessment, we reveal nuanced computational discrepancies that static graph features miss. Our experiments on instance generative models demonstrate that instances with high structural similarity scores can still exhibit drastically divergent solver interactions and difficulty levels. By providing this multifaceted evaluation toolkit, GenBench-MILP aims to facilitate rigorous comparisons and guide the development of high-fidelity instance generators.

MILP生成评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。