arXiv:2605.08305cs.LGcs.AI2026-05

首个面向真实大模型系统的超参优化基准,助力AutoML研究

LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems

论文配图:LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems
图 1 · 摘自论文原文
  • 构建涵盖12-23维超参、3-5维保真度的实时基准套件
  • 包含36万+配置,支持3-9个推理指标与2-10种成本度量
  • 为自动化机器学习提供可演进的研究平台,适合系统与优化研究者

大语言模型(LLM)系统在众多应用领域成为AI前沿,给自动化机器学习(AutoML)中的超参数优化(HPO)带来新挑战与机遇。然而,此类系统存在来自AI与非AI组件的复合超参空间、保真度因素带来的丰富非线性影响,以及测量配置的多样成本,现有基准尚未全面覆盖。本文提出首个(实时)基准套件与数据集——LLMSYS-HPOBench,涵盖从真实运行中采集的超参配置推理目标值。当前包含364,450条超参配置,维度为12-23,保真度维度3-5,对应932种设置,支持3-9个推理目标指标与2-10种成本度量,并附有测量生成的日志。我们不仅呼吁对现有HPO算法在前沿大模型系统上的再验证,更希望为AutoML社区提供一个持续演进的研究平台。基准已开源:https://github.com/ideas-labo/llmsys-hpobench

原文摘要 · Abstract (English)

Large Language Model (LLM) systems have been the frontier of AI in many application domains, leading to new challenges and opportunities for hyperparameter optimization (HPO) for the AutoML community. However, this type of system exhibits an unprecedented compound space of hyperparameter configuration from both the AI and non-AI components; rich and nonlinear implications from the fidelity factors; and diverse costs of measuring hyperparameter configurations, none of which have been fully captured in existing benchmarks. This paper presents the first (live) benchmark suite and datasets for HPO of real-world LLM systems, dubbed LLMSYS-HPOBench, covering data related to the inference objective values of hyperparameter configurations profiled from running the LLM systems. Currently, LLMSYS-HPOBench contains 364,450 hyperparameter configurations with a dimensionality of 12-23, 3-5 dimensions of fidelity factor leading to 932 settings, 3-9 inference objective metrics, and 2-10 cost metrics, together with generated logs from measuring the LLM systems. What we seek to advocate is not only a revalidation of the existing HPO algorithms over the frontier LLM systems, but also to provide an evolving platform for the AutoML community to explore new directions of research in this regard. The benchmark suite has been made available at: https://github.com/ideas-labo/llmsys-hpobench

超参优化大模型系统AutoML基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。