arXiv:2505.24541cs.CVcs.AI2025-05被引 4

用多专家动态路由解决多模态模型的视觉任务冲突问题

Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts

  • 设计多视觉专家混合架构,按任务动态分配图像处理
  • 在多个视觉任务上实现显著性能提升,计算开销极低
  • 可无缝嵌入任意多模态大模型,适合需要多任务协同的场景

多模态大语言模型(MLLMs)需对复杂图像信息进行细致理解,通常依赖视觉编码器感知不同视觉场景。但仅使用单一视觉编码器处理多样化任务域,易引发冲突。现有方法通过直接集成多个领域专用视觉编码器提升数据感知能力,但结构复杂且难以联合优化。本文提出 Mixpert,一种高效的视觉专家混合架构,在保持单编码器联合学习优势的同时,重构为多专家范式,支持跨视觉任务的特定微调。此外,设计动态路由机制,将输入图像分配给最合适的视觉专家。Mixpert有效缓解了单个视觉编码器在多任务学习中的领域冲突,且额外计算成本极小,优于多个编码器方案。实验表明,该方法可无缝集成至任意 MLLM,各项任务性能均有显著提升。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) require a nuanced interpretation of complex image information, typically leveraging a vision encoder to perceive various visual scenarios. However, relying solely on a single vision encoder to handle diverse task domains proves difficult and inevitably leads to conflicts. Recent work enhances data perception by directly integrating multiple domain-specific vision encoders, yet this structure adds complexity and limits the potential for joint optimization. In this paper, we introduce Mixpert, an efficient mixture-of-vision-experts architecture that inherits the joint learning advantages from a single vision encoder while being restructured into a multi-expert paradigm for task-specific fine-tuning across different visual tasks. Additionally, we design a dynamic routing mechanism that allocates input images to the most suitable visual expert. Mixpert effectively alleviates domain conflicts encountered by a single vision encoder in multi-task learning with minimal additional computational cost, making it more efficient than multiple encoders. Furthermore, Mixpert integrates seamlessly into any MLLM, with experimental results demonstrating substantial performance gains across various tasks.

多模态视觉专家动态路由高效架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。