arXiv:2512.14098cs.LGcs.DC2025-12被引 5

自动规划多模态模型部署,显著提升服务吞吐量。

Cornfigurator: Automated Planning for Any-to-Any Multimodal Model Serving

  • 基于模型与负载特征,自动探索全谱部署策略。
  • 相比现有系统和人工调优,吞吐量提升1.12到6.32倍。
  • 适合需要高效部署任意输入输出模态模型的工程师。

Any-to-Any 模型是一类新兴的多模态模型,可接受文本与多模态数据的任意组合作为输入并生成相应输出,带来异构计算路径和组件扩展特性。现有部署机制要么需手动调优,要么无法泛化至通用 Any-to-Any 模型。本文提出 Cornfigurator,首个针对通用 Any-to-Any 模型推理服务的部署规划器。目标是最大化整体 goodput(即满足延迟要求的请求吞吐量)。Cornfigurator 基于模型与工作负载特性,探索从共置到解耦及混合策略的完整部署空间,并通过粗粒度到细粒度的统计评估高效搜索。生成的部署方案在性能上优于或达到现有系统与专家调优方案,好比吞吐量提升1.12×至6.32×。

原文摘要 · Abstract (English)

Any-to-Any models are an emerging class of multimodal models that accept combinations of text and multimodal data as input and generate them as output, introducing heterogeneous computation paths and component scaling characteristics. There are existing mechanisms for deploying Any-to-Any models--or special cases of them--for inference serving, but they either require manual effort and expertise to tune, or do not generalize to generic Any-to-Any models. We present Cornfigurator, the first deployment planner for generic Any-to-Any model inference serving. The goal of Cornfigurator is to maximize the overall goodput of serving the model, defined as the throughput of requests meeting their latency targets. To do so, based on model and workload characteristics, Cornfigurator explores the full spectrum of deployment strategies, from colocation to disaggregation and mixing different strategies. Cornfigurator performs coarse-to-fine statistical evaluation to efficiently navigate the large space of candidate plans. Plans generated by Cornfigurator either match or deliver 1.12$\times$-6.32$\times$ higher goodput compared to existing systems and expert-tuned deployment plans.

多模态部署优化推理服务自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。