为联邦式AI即服务设计高保真网络管理框架,确保跨域服务稳定可靠。
High-Fidelity Network Management for Federated AI-as-a-Service: Cross-Domain Orchestration
- 用尾部风险包络建模各域性能,结合确定性与随机性约束
- 实现p99.9延迟达标率,支持负载过载下的准入控制
- 适用于需要强性能保障的AIaaS运营商和多租户系统
为支撑AI即服务(AIaaS)的发展,通信服务提供商(CSPs)正从单纯连接提供者转型为管理型AIaaS网络服务(具备控制与编排能力的平面),负责将用户意图转化为模型并实现端到端的网络-计算协同编排。在此模式下,服务可靠性由通信损耗(延迟、丢包)与推理性能(延迟、错误)共同决定。核心挑战在于构建跨域联邦场景下的高保真运营管理体系。本文提出基于尾部风险包络(Tail-Risk Envelopes, TREs)的保障导向管理平面:一种可组合的、带签名的域级描述符,融合确定性约束与随机性的速率-延迟-扰动模型。利用随机网络演算,推导出串联域间端到端延迟违规概率的上界,并获得可用于优化的风险预算分解。结果显示,租户级预留机制可防止突发流量导致尾部延迟升高。审计层通过运行时遥测数据估算极端分位数性能,量化不确定性,并将尾部风险归因至各域以实现问责。包级蒙特卡洛仿真表明,在过载条件下,通过准入控制可提升p99.9合规性;在相关突发流量下仍能保持稳健的租户隔离。
原文摘要 · Abstract (English)
To support the emergence of AI-as-a-Service (AIaaS), communication service providers (CSPs) are on the verge of a radical transformation-from pure connectivity providers to AIaaS a managed network service (control-and-orchestration plane that exposes AI models). In this model, the CSP is responsible not only for transport/communications, but also for intent-to-model resolution and joint network-compute orchestration, i.e., reliable and timely end-to-end delivery. The resulting end-to-end AIaaS service thus becomes governed by communications impairments (delay, loss) and inference impairments (latency, error). A central open problem is an operational AIaaS control-and-orchestration framework that enforces high fidelity, particularly under multi-domain federation. This paper introduces an assurance-oriented AIaaS management plane based on Tail-Risk Envelopes (TREs): signed, composable per-domain descriptors that combine deterministic guardrails with stochastic rate-latency-impairment models. Using stochastic network calculus, we derive bounds on end-to-end delay violation probabilities across tandem domains and obtain an optimization-ready risk-budget decomposition. We show that tenant-level reservations prevent bursty traffic from inflating tail latency under TRE contracts. An auditing layer then uses runtime telemetry to estimate extreme-percentile performance, quantify uncertainty, and attribute tail-risk to each domain for accountability. Packet-level Monte-Carlo simulations demonstrate improved p99.9 compliance under overload via admission control and robust tenant isolation under correlated burstiness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。