提出首个统一评估大数据批处理自动扩缩的开源框架
BatchBench: Toward a Workload-Aware Benchmark for Autoscaling Policies in Big Data Batch Processing -- A Proposed Framework
- 构建六类批处理工作负载分类体系
- 设计可参数化的工作负载生成器与验证方法
- 支持规则、学习与大模型策略统一评测
自动扩缩已成为云原生大数据处理的基本需求,其设计范式已从基于规则的启发式方法扩展至学习型控制器和大型语言模型(LLM)代理。然而,尽管相关研究不断增多,社区仍缺乏共享基准进行对比。现有评估依赖合成TPC类查询、厂商博客中的专有基线或狭窄的迹线重放,导致不同策略在不同工作负载、基线和成本模型下报告有利结果,难以跨论文比较。本文为立场论文,提出BatchBench——一个开放基准框架,旨在使规则、学习与智能体式自动扩缩策略在同等实验条件下评估。贡献包括:(1) 基于公开基准与集群迹线提炼出六类批处理工作负载分类;(2) 设计参数化工作负载生成器,并采用两样本Kolmogorov-Smirnov检验与地球移动距离进行验证;(3) 构建五轴评估框架,涵盖成本、服务等级协议达成率、扩缩响应性、扩缩震荡与决策可解释性,且首度纳入LLM推理成本;(4) 定义标准化智能体接口,实现基于LLM与强化学习的扩缩器与规则控制器通过单一API对比。本文讨论预期评估空间,识别待解研究问题,并规划后续实证论文路线图。BatchBench参考实现正在开发中,将开源发布。
原文摘要 · Abstract (English)
Autoscaling has become a baseline expectation for cloud-native big data processing, and the design space has expanded beyond rule-based heuristics to include learned controllers and, most recently, large language model (LLM) agents. Yet despite a growing body of work spanning these paradigms, the community lacks a shared benchmark for comparing them. Existing evaluations rely on synthetic TPC-style queries, vendor blog posts with proprietary baselines, or narrow trace replays. Each new policy reports favorable numbers against a different baseline, on a different workload, with a different cost model, making cross-paper comparison effectively impossible. This is a position paper. We propose BatchBench, an open benchmarking framework designed to place rule-based, learned, and agentic autoscaling policies on equal experimental footing. The contribution is the design of the framework, not empirical results. We contribute: (1) a workload taxonomy of six batch processing classes synthesized from published autoscaling benchmarks and publicly released cluster traces; (2) the design of a parameterized workload generator with a validation methodology based on two-sample Kolmogorov-Smirnov and earth-mover distance; (3) a five-axis evaluation harness specification covering cost, SLA attainment, scaling responsiveness, scaling thrash, and decision interpretability, with first-class accounting for LLM inference cost; and (4) a standardized agent interface that lets LLM-based and reinforcement-learning autoscalers be evaluated alongside rule-based controllers with a single API. We discuss the expected evaluation surface, identify open research questions the framework is designed to answer, and outline a roadmap for the empirical paper that will follow. BatchBench's reference implementation is in active development and will be released as open source.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。