arXiv:2605.23417cs.LG2026-05

首个开源黑箱优化数据集,助力大模型学习通用优化方法

An Open-Source Training Dataset for Foundation Models for Black-box Optimization

论文配图:An Open-Source Training Dataset for Foundation Models for Black-box Optimization
图 1 · 摘自论文原文
  • 构建包含50万条轨迹的公开数据集,覆盖3095个真实黑箱问题
  • 训练200万至8000万参数模型,验证大规模预训练有效性
  • 适合研究优化算法、机器学习可迁移性的研究人员

大多数黑箱优化方法需大量超参数调优,常限制其在不同优化领域中的泛化能力。基于大规模优化轨迹学习优化原理的优化基础模型提供了一种有前景的替代方案,有望在多种问题类别中超越人工设计的方法。然而,先前工作或依赖非公开数据集,或仅使用纯合成数据,限制了可复现性和对现实问题的泛化能力。因此,该领域进展受限于缺乏大规模、真实世界、公开可用的预训练数据。我们提出BBO-Pile,首个开源数据集,包含超过50万条优化轨迹,覆盖3095个不同黑箱问题,适用于多种优化器,是目前该任务下最大的公开数据集。利用此数据集,我们训练了参数量从200万到8000万、训练样本量从2亿到20亿词元的多尺度基础模型,并研究其与计算资源的缩放关系。结果表明,大规模预训练是模拟黑箱优化方法的可行且有效路径,为未来研究铺平道路。

原文摘要 · Abstract (English)

Most black-box optimization methods require extensive hyperparameter tuning, often limiting their ability to generalize across different optimization domains. Foundation models for black-box optimization that learn optimization principles from a large collection of optimization trajectories offer a promising alternative, with the potential to outperform manually designed methods across diverse problem classes. However, prior work has either relied on non-public datasets or on purely synthetic data, limiting reproducibility and generalization to real-world problems. As a result, progress in this area has been constrained by the lack of large-scale, real-world, publicly available pre-training data. We introduce BBO-Pile, the first open-source dataset comprising over 500K optimization trajectories evaluated across 3095 different black-boxes for different optimizers, which represents by far the largest public dataset for this task. Using this dataset, we train a family of foundation models at multiple scales, ranging from 2M to 80M parameters and from 200M to 2B training tokens, and study their scaling behavior with respect to compute. Our results demonstrate that large-scale pre-training is a viable and effective approach to imitate black-box optimization methods, paving the way for future research in this direction.

黑箱优化基础模型数据集可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。