用多模型优势互补,提升多任务视觉学习性能
Swiss Army Knife: Synergizing Biases in Knowledge from Vision Foundation Models for Multi-Task Learning
- 通过自适应融合多个视觉模型的特征表示
- 在NYUD-v2上比现有方法提升10%准确率
- 适合需要多任务协同的视觉系统开发
视觉基础模型(VFMs)在众多下游任务中表现优异,但因训练范式差异导致其在不同任务中存在固有表征偏差。尽管融合多个VFMs的优势是直观思路,但有效利用这些偏差仍是挑战。本文提出一种灵活的“瑞士军刀”(SAK)方案,通过轻量级教师专用适配路径与通用骨干结构协同,动态选择并组合多模型表示。采用混合表示路由机制,实现多个视觉基础模型优势的互补。大量实验表明,SAK在多任务学习中显著优于现有最优方法,在NYUD-v2基准上提升10%,同时具备良好可扩展性与鲁棒性,支持未来模型升级。
原文摘要 · Abstract (English)
Vision Foundation Models (VFMs) have demonstrated outstanding performance on numerous downstream tasks. However, due to their inherent representation biases originating from different training paradigms, VFMs exhibit advantages and disadvantages across distinct vision tasks. Although amalgamating the strengths of multiple VFMs for downstream tasks is an intuitive strategy, effectively exploiting these biases remains a significant challenge. In this paper, we propose a novel and versatile "Swiss Army Knife" (SAK) solution, which adaptively distills knowledge from a committee of VFMs to enhance multi-task learning. Unlike existing methods that use a single backbone for knowledge transfer, our approach preserves the unique representation bias of each teacher by collaborating the lightweight Teacher-Specific Adapter Path modules with the Teacher-Agnostic Stem. Through dynamic selection and combination of representations with Mixture-of-Representations Routers, our SAK is capable of synergizing the complementary strengths of multiple VFMs. Extensive experiments show that our SAK remarkably outperforms prior state of the arts in multi-task learning by 10% on the NYUD-v2 benchmark, while also providing a flexible and robust framework that can readily accommodate more advanced model designs. Project page: https://innovator-zero.github.io/SAK/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。