arXiv:2606.03444cs.CVcs.AI2026-06中稿 · ICML

通过专家自组织分工,融合多视觉模型优势,提升效率与精度。

PRISM: Synergizing Vision Foundation Models via Self-organized Expert Specialization

论文配图:PRISM: Synergizing Vision Foundation Models via Self-organized Expert Specialization
图 1 · 摘自论文原文
  • 采用双流MoE架构,让专家模块专注不同表征空间,减少干扰。
  • 在PASCAL-Context和NYUD-v2上达到新最优性能,验证可扩展性。
  • 适合需要高效集成多视觉模型的下游任务开发者使用。

将多种视觉基础模型(VFMs)的互补优势整合到单一高效模型中极具价值,但受制于单体蒸馏带来的负迁移问题。为解决特征冲突,我们提出新型双流混合专家(MoE)框架PRISM,通过模块化专长实现视觉模型协同。采用两阶段范式:(1) 专长解构,教师条件路由引导专家聚焦于不同表征子空间以减轻干扰;(2) 动态重组,路由学习为下游任务定制计算路径。在PASCAL-Context和NYUD-v2上的实验表明,PRISM建立了新基准,验证了稀疏、涌现的专长是一种可扩展的多样化视觉知识整合方法。

原文摘要 · Abstract (English)

Unifying the complementary strengths of diverse Vision Foundation Models (VFMs) into a single efficient model is highly desirable but challenged by the negative transfer inherent in monolithic distillation. To address these feature conflicts, we introduce \textbf{PRISM}, a novel dual-stream Mixture-of-Experts (MoE) framework that synergizes VFMs via modular specialization. We propose a two-stage paradigm: (1) expertise deconstruction, where a teacher-conditional router guides experts to specialize in distinct representational subspaces to mitigate interference, followed by (2) dynamic recomposition, where the router learns to assemble these experts into tailored computational pathways for downstream tasks. Experiments on PASCAL-Context and NYUD-v2 show that \textbf{PRISM} establishes a new state of the art, validating that sparse, emergent specialization is a scalable approach for integrating diverse visual knowledge.

视觉模型MoE专家系统知识融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。