arXiv:2603.06403cs.LG2026-03被引 2

针对多模态大模型调度难题,提出自适应增强的在线决策框架。

Adapter-Augmented Bandits for Online Multi-Constrained Multi-Modal Inference Scheduling

  • 用轻量适配器提取任务特征,实现高效在线更新
  • 在多维预算约束下,最高提升14.18%任务奖励
  • 适合资源受限的实时多模态推理系统部署

多模态大语言模型(MLLM)推理调度可在实际异构资源预算下实现优异响应质量,超越单一后端设置的能力。然而在线调度极具挑战:请求在模态组合与隐式推理难度上差异显著,而执行后端因系统抖动和网络波动导致成本动态变化。二者耦合不确定性带来两大核心难题:生成语义准确且调度相关的多模态任务表示,以及在不可逆的多维预算下进行低开销在线决策。为此,我们提出 M-CMAB(多模态多约束上下文多臂老虎机),包含三个组件:(i) 基于CLS注意力、冻结主干的预测器,仅更新轻量适配器以完成动作特定估计;(ii) 原始-对偶约束器,通过每轮目标维护在线拉格朗日乘子,实现长周期约束;(iii) 两阶段调度器,在不可逆预算下平衡探索与利用。我们在含异构后端的复合多模态基准上建立了多维背包约束下的后悔率保证。M-CMAB 在各类预算条件下持续优于现有最优基线,最高达14.18%奖励提升,并紧密逼近基于理想信息的上界。代码已公开于 https://anonymous.4open.science/r/M2CMAB/。

原文摘要 · Abstract (English)

Multi-modal large language model (MLLM) inference scheduling enables strong response quality under practical and heterogeneous budgets, beyond what a homogeneous single-backend setting can offer. Yet online MLLM task scheduling is nontrivial, as requests vary sharply in modality composition and latent reasoning difficulty, while execution backends incur distinct, time-varying costs due to system jitter and network variation. These coupled uncertainties pose two core challenges: deriving semantically faithful yet scheduling-relevant multi-modal task representations, and making low-overhead online decisions over irreversible multi-dimensional budgets. Accordingly, we propose \emph{M-CMAB} (\underline{M}ulti-modal \underline{M}ulti-constraint \underline{C}ontextual \underline{M}ulti-\underline{A}rmed \underline{B}andit), a multi-adapter-enhanced MLLM inference scheduling framework with three components: (i) a CLS-attentive, frozen-backbone \emph{Predictor} that extracts compact task representations and updates only lightweight adapters for action-specific estimation; (ii) a primal-dual \emph{Constrainer} that maintains online Lagrange multipliers to enforce long-horizon constraints via per-round objectives; and (iii) a two-phase \emph{Scheduler} that balances exploration and exploitation under irreversible budgets. We establish a regret guarantee under multi-dimensional knapsack constraints. On a composite multimodal benchmark with heterogeneous backends, \emph{M-CMAB} consistently outperforms state-of-the-art baselines across budget regimes, achieving up to 14.18% higher reward and closely tracking an oracle-aided upper bound. Codes are available at https://anonymous.4open.science/r/M2CMAB/.

多模态在线调度强化学习资源管理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。