多模态边缘计算中,压缩、路由与量化相互影响,需协同优化。
Beyond Independent Optimization: Compression, MoE Routing, and Quantization Interactions in Multimodal Edge Intelligence

- 揭示压缩、路由、量化三者间的动态交互机制
- 发现量化误差会放大路由偏差,影响专家利用率
- 适合研究多模态模型部署与边缘计算的工程师
高效多模态推理不仅受限于模型质量或计算量,更受延迟、内存和能耗约束下表征保存、传输、路由、缓存与量化成本的影响。本文综述视觉-语言及多模态大模型在视觉标记压缩、视频标记管理、KV-cache优化、混合专家(MoE)路由、低比特量化、边缘部署与硬件感知基准测试方面的进展。指出这些技术不可独立优化:视觉压缩改变下游特征分布并影响MoE路由决策,路由行为决定专家利用率与量化敏感度,量化后的路由器逻辑影响专家分配,缓存策略决定保留的多模态证据,硬件约束常使计算节省转化为内存与通信瓶颈。我们围绕这些交互组织文献,识别关键权衡,包括精度与标记预算、静态与自适应压缩、稀疏路由效率与专家崩溃、低比特推理与模态特异性退化。最后引入时序路由一致性作为视频MoE模型诊断工具,并提出路由感知压缩、跨模态缓存管理、硬件感知协同设计与统一基准评测等开放方向。
原文摘要 · Abstract (English)
Efficient multimodal inference is increasingly constrained not only by model quality or FLOP count, but also by the cost of preserving, moving, routing, caching, and quantizing multimodal representations under latency, memory, and energy constraints. This paper reviews recent advances in efficient vision-language and multimodal large language models, covering visual token compression, video token management, KV-cache optimization, Mixture-of-Experts (MoE) routing, low-bit quantization, edge deployment, and hardware-aware benchmarking. We argue that these techniques cannot be treated as independent optimizations. Visual token compression alters downstream feature distributions and MoE routing decisions, routing behavior affects expert utilization and quantization sensitivity, quantized router logits influence expert assignment, KV-cache policies determine retained multimodal evidence, and hardware constraints often transform computational savings into memory and communication bottlenecks. We organize the literature around these interactions and identify key design trade-offs, including accuracy versus token budget, static versus adaptive compression, sparse routing efficiency versus expert collapse, and low-bit inference versus modality-specific degradation. Finally, we introduce Temporal Routing Consistency as a diagnostic for video MoE models and highlight open research directions in routing-aware compression, cross-modal cache management, hardware-aware co-design, and unified benchmarking for multimodal edge intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。