arXiv:2603.24084cs.AI2026-03

首个标准化多目标搜索基准,解决评估不统一问题。

Bridging the Evaluation Gap: Standardized Benchmarks for Multi-Objective Search

  • 构建跨四类场景的统一基准,包含固定图结构和标准查询
  • 覆盖强相关到完全独立的目标关系,支持精确与近似评估
  • 适合算法开发者、评估研究者及机器人路径规划领域

多目标搜索(MOS)的实证评估长期存在碎片化问题,依赖异构问题实例和不可比的目标定义,导致跨研究比较困难。尤其值得注意的是,传统默认基准DIMACS道路网络的目标高度相关,无法体现多样化的帕累托前沿结构。为此,我们提出首个全面且标准化的精确与近似多目标搜索基准套件。该套件涵盖四类结构迥异的领域:真实世界道路网络、结构化合成图、基于游戏的网格环境以及高维机器人运动规划路线图。通过提供固定的图实例、标准的起止点查询、参考的精确帕累托最优解集,以及标准化的近似(ε-支配)评估协议,该套件完整捕捉了目标间的多种交互模式——从高度相关到严格独立。最终,该基准为未来MOS评估提供了可复现、鲁棒且结构全面的共同基础。

原文摘要 · Abstract (English)

Empirical evaluation in multi-objective search (MOS) has historically suffered from fragmentation, relying on heterogeneous problem instances with incompatible objective definitions that make cross-study comparisons difficult. This standardization gap is further exacerbated by the realization that DIMACS road networks, a historical default benchmark for the field, exhibit highly correlated objectives that fail to capture diverse Pareto-front structures. To address this, we introduce the first comprehensive, standardized benchmark suite for exact and approximate MOS. Our suite spans four structurally diverse domains: real-world road networks, structured synthetic graphs, game-based grid environments, and high-dimensional robotic motion-planning roadmaps. By providing fixed graph instances, standardized start-goal queries, reference exact Pareto-optimal solution sets, and a standardized approximate ($\varepsilon$-dominance) evaluation protocol, this suite captures a full spectrum of objective interactions: from strongly correlated to strictly independent. Ultimately, this benchmark provides a common foundation to ensure future MOS evaluations are robust, reproducible, and structurally comprehensive.

多目标搜索基准测试路径规划算法评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。