用生成式马尔可夫模型优化分布式系统资源调度
Brief Announcement: Generative Markov Model for Distributed Computing Systems
- 将分布式系统建模为结构化状态的生成式马尔可夫模型
- 分布式计算可降低延迟并减少服务器资源消耗
- 适合研究资源调度与强化学习在边缘计算中的应用
新兴的分布式计算范式(如计算连续体)具有异构性、随机性和复杂性。高效利用全链路资源需要统一的系统形式化模型。为此,我们提出一种通用框架,将分布式计算系统建模为在结构化系统状态上因子分解的生成式马尔可夫模型。该模型将状态分解为高维变量,每个变量进一步在其元素上因子分解,反映分布式系统固有的稀疏依赖结构。这一建模方法使原本难以处理的系统状态得以实现模拟、推断和策略学习,连接了分布式计算与马尔可夫链理论及强化学习(RL)。通过协同AI推理的案例研究,我们发现集中式调度在大规模下成为瓶颈,而将计算分布到用户设备可降低延迟并减少服务器资源消耗。这些结果凸显自适应决策的价值,并验证了该框架在建模、仿真与优化方面的实用性。
原文摘要 · Abstract (English)
Emerging distributed computing paradigms, such as the computing continuum, are inherently heterogeneous, stochastic, and complex. Efficiently and effectively utilizing all available resources across the continuum demands a unified formal model of the system. To address this gap, we propose a general framework for modeling distributed computing systems as a generative Markov model, factorized over a structured system state. In our model, the state decomposes into high-dimensional variables, each further factorized over its elements, reflecting the sparse dependency structure inherent to distributed systems. This yields a tractable model enabling simulation, inference, and policy learning over otherwise intractable system states, bridging distributed computing with Markov chain theory and reinforcement learning (RL). We demonstrate our framework through a case study of collaborative AI inference, in which a dedicated server combines resources with those volunteered by service users. Our results show that centralized scheduling becomes a bottleneck at scale, while distributing computation across user devices reduces both latency and server resource consumption. These findings highlight the value of adaptive decision-making in distributed computing systems and demonstrate the framework's utility for modeling, simulation, and optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。