arXiv:2512.11588cs.AI2025-12被引 2

让AI评估更公平实用,推动动态、可复现的 benchmark 普及。

AI Benchmark Democratization and Carpentry

  • 构建动态适应的评估框架,随模型和数据演进。
  • 强调真实场景应用相关性,而非仅追求顶尖硬件峰值性能。
  • 倡导教育与社区协作,培养全民可用的 benchmark 设计能力。

基准测试是现代机器学习的核心,促进可复现性、对比和科学进步。然而,AI基准测试日益复杂,需动态、面向AI的工作流。模型架构、规模、数据集和部署环境的快速演进使评估成为不断变化的目标。大语言模型常记忆静态基准,导致基准结果与实际性能脱节。除传统静态基准外,亟需持续自适应的评估框架,以对齐科学评估与部署风险。这需要掌握‘基准搭建’技能与教育。基于MLCommons、教育项目及美国能源部万亿参数联盟的经验,主要障碍包括高资源需求、专用硬件访问受限、基准设计专业知识缺乏,以及结果与应用场景关联性不明确。当前基准多聚焦顶级硬件上的峰值性能,对多样化真实场景指导有限。基准测试必须动态化,融合演进中的模型、更新的数据和异构平台,同时保持透明、可复现、可解释。普及化需技术革新与系统性教育并行,建立可持续的基准设计与使用专业能力。基准应支持应用相关的比较,助力做出知情的、情境敏感的决策。动态、包容的基准测试将确保评估跟上AI发展步伐,支撑负责任、可复现、可访问的AI部署。社区努力可为‘基准搭建’奠定基础。

原文摘要 · Abstract (English)

Benchmarks are a cornerstone of modern machine learning, enabling reproducibility, comparison, and scientific progress. However, AI benchmarks are increasingly complex, requiring dynamic, AI-focused workflows. Rapid evolution in model architectures, scale, datasets, and deployment contexts makes evaluation a moving target. Large language models often memorize static benchmarks, causing a gap between benchmark results and real-world performance. Beyond traditional static benchmarks, continuous adaptive benchmarking frameworks are needed to align scientific assessment with deployment risks. This calls for skills and education in AI Benchmark Carpentry. From our experience with MLCommons, educational initiatives, and programs like the DOE's Trillion Parameter Consortium, key barriers include high resource demands, limited access to specialized hardware, lack of benchmark design expertise, and uncertainty in relating results to application domains. Current benchmarks often emphasize peak performance on top-tier hardware, offering limited guidance for diverse, real-world scenarios. Benchmarking must become dynamic, incorporating evolving models, updated data, and heterogeneous platforms while maintaining transparency, reproducibility, and interpretability. Democratization requires both technical innovation and systematic education across levels, building sustained expertise in benchmark design and use. Benchmarks should support application-relevant comparisons, enabling informed, context-sensitive decisions. Dynamic, inclusive benchmarking will ensure evaluation keeps pace with AI evolution and supports responsible, reproducible, and accessible AI deployment. Community efforts can provide a foundation for AI Benchmark Carpentry.

基准测试AI评估开源生态教育普及

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。