arXiv:2608.11840cs.DCcs.AI2026-08

让用户共享算力,降低AI推理服务成本。

User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling

论文配图:User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling
图 1 · 摘自论文原文
  • 用用户自愿贡献的资源协同处理任务,减少对中心化资源依赖。
  • 用户越多,分布式调度越高效,请求完成率提升,延迟降低。
  • 适合大规模AI服务部署者参考,尤其关注成本与性能平衡的场景。

随着人工智能推理服务需求增长,集中式服务基础设施成本随之上升。本文提出一种结合专用资源与用户自愿贡献资源的协同分布式推理系统。专用资源保障服务质量(QoS),而用户贡献的资源则应对需求增长,避免中心化设施按比例扩容。为建模用户、资源、任务与策略间的随机动态交互,我们构建了具有结构化时间因子分解的高维生成马尔可夫模型,支持仿真,并为任务调度与QoS感知的资源分配优化提供基础。我们在不同用户规模、资源容量及集中/分布式调度策略下进行评估。仿真结果表明,随着用户数量增加,分布式调度优势愈发显著,有效提升请求完成率和降低P99延迟,同时大幅减少专用资源消耗。这证明了用户协同推理在实现高效自动伸缩方面的可行性。

原文摘要 · Abstract (English)

Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise with demand. We propose a collaborative distributed inference system combining dedicated infrastructure with resources contributed by service users. Dedicated resources provide baseline capacity for maintaining quality of service (QoS), while volunteered resources absorb increasing demand without proportional growth in centralized infrastructure. To capture stochastic and dynamic interactions among users, resources, tasks, and policies, we develop a high-dimensional generative Markov model with structured temporal factorization. The model supports simulation and provides a foundation for task scheduling and QoS-aware resource allocation optimization. We evaluate the system across user populations, resource capacities, and centralized and distributed scheduling policies. Simulations show that distributed scheduling becomes increasingly advantageous as the user population grows, improving request completion and P99 latency while substantially reducing dedicated resource consumption. These results demonstrate the feasibility of user-assisted collaborative inference for infrastructure-efficient autoscaling.

分布式推理资源协同自动伸缩QoS优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。