arXiv:2602.22938cs.CVcs.AI2026-02ICLR被引 1

用多个专家提示词提升视觉适应任务表现

pMoE: Prompting Diverse Experts Together Wins More in Visual Adaptation

  • 引入专家专用提示词与可学习调度器,动态整合多领域知识
  • 在47个任务上显著优于现有方法,计算效率与适应性平衡更优
  • 适合需要跨领域适应的视觉模型微调场景

参数高效微调在各类视觉适应任务(如分类与分割)中展现出良好效果。传统提示微调通常仅利用单一预训练模型的知识,无论通用或医学领域。然而,该方法常忽略在同一微调过程中融合多元领域知识的潜力。本文提出一种新型混合专家提示微调方法pMoE,通过专家专属提示词与可学习调度器,有效在统一模型框架中整合多个专家领域的优势。pMoE在不同提示层采用动态令牌调度机制,优化各领域专家在适应阶段的贡献。通过融合多样化领域知识,所提方法显著提升模型的泛化能力与任务适用范围。我们在47个适应任务上进行了广泛实验,涵盖通用与医学领域的分类和分割任务。结果表明,pMoE不仅性能显著领先,且在计算效率与适应有效性之间实现更优权衡。

原文摘要 · Abstract (English)

Parameter-efficient fine-tuning has demonstrated promising results across various visual adaptation tasks, such as classification and segmentation. Typically, prompt tuning techniques have harnessed knowledge from a single pre-trained model, whether from a general or a specialized medical domain. However, this approach typically overlooks the potential synergies that could arise from integrating diverse domain knowledge within the same tuning process. In this work, we propose a novel Mixture-of-Experts prompt tuning method called pMoE, which leverages the strengths of multiple expert domains through expert-specialized prompt tokens and the learnable dispatcher, effectively combining their expertise in a unified model framework. Our pMoE introduces expert-specific prompt tokens and utilizes a dynamic token dispatching mechanism at various prompt layers to optimize the contribution of each domain expert during the adaptation phase. By incorporating both domain knowledge from diverse experts, the proposed pMoE significantly enhances the model's versatility and applicability to a broad spectrum of tasks. We conduct extensive experiments across 47 adaptation tasks, including both classification and segmentation in general and medical domains. The results demonstrate that our pMoE not only achieves superior performance with a large margin of improvements but also offers an optimal trade-off between computational efficiency and adaptation effectiveness compared to existing methods.

提示微调多专家视觉适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。