提出可上传的多源少样本域适应框架,降低边缘设备负担。
Vision-aware Multimodal Prompt Tuning for Uploadable Multi-source Few-shot Domain Adaptation
- 设计视觉感知的跨模态提示调优方法,提升语义区分度。
- 在OfficeHome和DomainNet上性能超越已有提示调优方法。
- 适合资源受限的边缘协同学习场景,仅需少量标注数据。
传统多源域少样本适应(MFDA)在低资源场景下难以进一步降低边缘设备负载。本文提出一种可上传的多源少样本域适应(UMFDA)框架,属于边缘侧模型去中心化协同学习,要求边缘模型保持低计算开销。源域数据中仅有少量标注,大部分为无标签数据。为此,本文提出在去中心化架构下的视觉感知跨模态提示调优框架(VAMP),其中视觉感知提示引导文本领域特定提示以维持语义判别性并感知域信息。通过跨模态语义与域分布对齐损失优化各边缘模型,同时利用文本分类一致性与语义多样性损失促进边缘模型间的协作学习。在OfficeHome和DomainNet数据集上的大量实验表明,所提VAMP在UMFDA中有效,性能优于先前提示调优方法。
原文摘要 · Abstract (English)
Conventional multi-source domain few-shot adaptation (MFDA) faces the challenge of further reducing the load on edge-side devices in low-resource scenarios. Considering the native language-supervised advantage of CLIP and the plug-and-play nature of prompt to transfer CLIP efficiently, this paper introduces an uploadable multi-source few-shot domain adaptation (UMFDA) schema. It belongs to a decentralized edge collaborative learning in the edge-side models that must maintain a low computational load. And only a limited amount of annotations in source domain data is provided, with most of the data being unannotated. Further, this paper proposes a vision-aware multimodal prompt tuning framework (VAMP) under the decentralized schema, where the vision-aware prompt guides the text domain-specific prompt to maintain semantic discriminability and perceive the domain information. The cross-modal semantic and domain distribution alignment losses optimize each edge-side model, while text classifier consistency and semantic diversity losses promote collaborative learning among edge-side models. Extensive experiments were conducted on OfficeHome and DomainNet datasets to demonstrate the effectiveness of the proposed VAMP in the UMFDA, which outperformed the previous prompt tuning methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。