用强化学习与神经网络实现云上AI推理的实时负载均衡与自动扩缩容。
Scalability Optimization in Cloud-Based AI Inference Services: Strategies for Real-Time Load Balancing and Automated Scaling
- 融合强化学习与深度神经网络,动态分配负载并预测需求。
- 负载均衡效率提升35%,响应延迟降低28%。
- 去中心化决策提升系统容错性,适合高并发云服务场景。
云上AI推理服务的快速扩张亟需稳健的可扩展性解决方案以应对动态负载并保持高性能。本文提出一个面向云AI推理服务的综合性可扩展性优化框架,重点聚焦实时负载均衡与自动扩缩容策略。所提模型采用混合方法,结合强化学习实现自适应负载分发,以及深度神经网络进行精准需求预测。该多层架构使系统能够预判工作负载波动并主动调整资源,从而最大化资源利用率并最小化延迟。此外,模型中引入去中心化决策机制,提升了故障容错能力,并缩短了扩容响应时间。实验结果表明,相比传统可扩展性方案,该模型负载均衡效率提升35%,响应延迟降低28%,展现出显著优化效果。
原文摘要 · Abstract (English)
The rapid expansion of AI inference services in the cloud necessitates a robust scalability solution to manage dynamic workloads and maintain high performance. This study proposes a comprehensive scalability optimization framework for cloud AI inference services, focusing on real-time load balancing and autoscaling strategies. The proposed model is a hybrid approach that combines reinforcement learning for adaptive load distribution and deep neural networks for accurate demand forecasting. This multi-layered approach enables the system to anticipate workload fluctuations and proactively adjust resources, ensuring maximum resource utilisation and minimising latency. Furthermore, the incorporation of a decentralised decision-making process within the model serves to enhance fault tolerance and reduce response time in scaling operations. Experimental results demonstrate that the proposed model enhances load balancing efficiency by 35\ and reduces response delay by 28\, thereby exhibiting a substantial optimization effect in comparison with conventional scalability solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。