arXiv:2506.19578cs.DCcs.AI2025-06被引 3

构建全球计算基础设施的动态模型,提升数据与任务分配效率。

Towards an Introspective Dynamic Model of Globally Distributed Computing Infrastructures

  • 基于PanDA系统五个月作业记录,提取排队时间等关键指标。
  • 构建生成式AI模型模拟任务负载,包含可见与不可见特征。
  • 适合研究分布式计算优化与智能调度的学者参考。

大型科学合作项目如ATLAS、Belle II、CMS、DUNE等涉及数百个研究机构和数千名研究人员,遍布全球。这些实验产生海量数据,规模即将达到艾字节级别。因此,对计算资源的需求急剧增长,涵盖从原始数据到可用衍生数据的结构化处理、大规模蒙特卡洛模拟以及各类终端用户分析。为应对计算与存储挑战,已部署集中式工作流与数据管理系统。然而,数据放置与任务分配决策往往独立进行且依赖启发式方法。采用更高效启发式或人工智能驱动方案的主要障碍在于缺乏快速可靠的动态模型来评估与优化替代策略。本研究旨在利用真实世界数据构建此类交互式系统。通过对PanDA工作流管理系统的作业执行记录进行分析,我们识别出排队时间、错误率及远程数据访问程度等关键性能指标。数据集涵盖五个月的活动记录。此外,我们正在开发一个生成式AI模型,用于模拟任务负载的时间序列,包含类别、事件数、提交组等显性特征,以及由历史PanDA记录和计算站点能力推导出的隐性特征(如总计算负载)。这些对任务分配器不可见的隐藏特征会影响排队时间和数据传输。

原文摘要 · Abstract (English)

Large-scale scientific collaborations like ATLAS, Belle II, CMS, DUNE, and others involve hundreds of research institutes and thousands of researchers spread across the globe. These experiments generate petabytes of data, with volumes soon expected to reach exabytes. Consequently, there is a growing need for computation, including structured data processing from raw data to consumer-ready derived data, extensive Monte Carlo simulation campaigns, and a wide range of end-user analysis. To manage these computational and storage demands, centralized workflow and data management systems are implemented. However, decisions regarding data placement and payload allocation are often made disjointly and via heuristic means. A significant obstacle in adopting more effective heuristic or AI-driven solutions is the absence of a quick and reliable introspective dynamic model to evaluate and refine alternative approaches. In this study, we aim to develop such an interactive system using real-world data. By examining job execution records from the PanDA workflow management system, we have pinpointed key performance indicators such as queuing time, error rate, and the extent of remote data access. The dataset includes five months of activity. Additionally, we are creating a generative AI model to simulate time series of payloads, which incorporate visible features like category, event count, and submitting group, as well as hidden features like the total computational load-derived from existing PanDA records and computing site capabilities. These hidden features, which are not visible to job allocators, whether heuristic or AI-driven, influence factors such as queuing times and data movement.

分布式计算生成模型任务调度科学计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。