动态调整多模态模型层数,按输入质量分配计算资源
ADMN: A Layer-Wise Adaptive Multimodal Network for Dynamic Input Noise and Compute Resources
- 按输入质量与算力约束,动态分配各模态的计算层数
- 在保持高精度的同时,降低75%浮点运算量
- 适合算力不稳或传感器易受干扰的实时场景
多模态深度学习系统因多种传感模态具备鲁棒性,被部署于动态场景。然而,它们常面临算力波动(如多租户、设备异构)和输入质量变化(如传感器噪声、环境干扰)的挑战。静态配置的多模态系统无法随算力变化自适应,现有动态网络又难以满足严格算力预算,且普遍忽略模态质量差异。导致严重受损的模态仍消耗大量资源,挤占本可用于优质模态的算力。为此,我们提出ADMN——一种逐层自适应深度多模态网络,可依据算力约束动态调节所有模态的活跃层数,并根据各模态输入质量持续重新分配计算资源。评估表明,ADMN在保持顶尖模型精度的同时,可减少高达75%的浮点运算量。
原文摘要 · Abstract (English)
Multimodal deep learning systems are deployed in dynamic scenarios due to the robustness afforded by multiple sensing modalities. Nevertheless, they struggle with varying compute resource availability (due to multi-tenancy, device heterogeneity, etc.) and fluctuating quality of inputs (from sensor feed corruption, environmental noise, etc.). Statically provisioned multimodal systems cannot adapt when compute resources change over time, while existing dynamic networks struggle with strict compute budgets. Additionally, both systems often neglect the impact of variations in modality quality. Consequently, modalities suffering substantial corruption may needlessly consume resources better allocated towards other modalities. We propose ADMN, a layer-wise Adaptive Depth Multimodal Network capable of tackling both challenges: it adjusts the total number of active layers across all modalities to meet strict compute resource constraints and continually reallocates layers across input modalities according to their modality quality. Our evaluations showcase ADMN can match the accuracy of state-of-the-art networks while reducing up to 75% of their floating-point operations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。