arXiv:2509.11413cs.LGcs.AI2025-09被引 1

将基准测试视为AI任务,实现动态优化与部署决策支持

Framing AI System Benchmarking as a Learning Task: FlexBench and the Open MLPerf Dataset

  • 把基准测试变成可学习的AI任务,持续评估多维度性能
  • 验证了在消费级服务器上对DeepSeek R1和LLaMA 3.3的高效评测能力
  • 适合需要成本优化部署的工程师与系统设计者

现有AI系统基准测试如MLPerf难以跟上快速演进的AI生态,难以为部署、优化与协同设计提供有效支持。本文提出将基准测试本身视为一个AI任务:模型在多样化数据集、软件和硬件上持续评估并优化,使用准确率、延迟、吞吐量、能耗和成本等关键指标。为此,我们推出了FlexBench——MLPerf LLM推理基准的模块化扩展,集成HuggingFace,可生成相关且可操作的洞察。评测结果与元数据被收集至Open MLPerf Dataset,支持协作维护、扩展,并用于预测建模与特征工程。通过MLPerf Inference提交成功验证了该理念,涵盖DeepSeek R1与LLaMA 3.3在通用服务器上的评估。目标是帮助从业者基于资源、需求与约束做出更经济高效的AI部署决策。

原文摘要 · Abstract (English)

Existing AI system benchmarks such as MLPerf often struggle to keep pace with the rapidly evolving AI landscape, making it difficult to support informed deployment, optimization, and co-design decisions for AI systems. We suggest that benchmarking itself can be framed as an AI task - one in which models are continuously evaluated and optimized across diverse datasets, software, and hardware, using key metrics such as accuracy, latency, throughput, energy consumption, and cost. To support this perspective, we present FlexBench: a modular extension of the MLPerf LLM inference benchmark, integrated with HuggingFace and designed to provide relevant and actionable insights. Benchmarking results and metadata are collected into an Open MLPerf Dataset, which can be collaboratively curated, extended, and leveraged for predictive modeling and feature engineering. We successfully validated the FlexBench concept through MLPerf Inference submissions, including evaluations of DeepSeek R1 and LLaMA 3.3 on commodity servers. The broader objective is to enable practitioners to make cost-effective AI deployment decisions that reflect their available resources, requirements, and constraints.

基准测试AI部署MLPerf动态优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。