提出多模态协作低秩分解框架,提升少样本视频域适应性能
Modality-Collaborative Low-Rank Decomposers for Few-Shot Video Domain Adaptation
- 将每种模态的特征按领域偏移程度分解为独有与共享部分
- 在三个公开数据集上显著超越现有方法,最高提升12.3%
- 适合研究少样本视频迁移学习与多模态特征融合的学者
本文研究少样本视频域适应(FSVDA)这一挑战性任务。视频的多模态特性引入独特难题,需同时考虑领域对齐与模态协作,而此前工作忽略此点。我们发现,在领域偏移影响下,单个模态及融合后多模态特征的泛化性能均受限,因各模态包含具有不同领域偏移特性的耦合特征,增加适应复杂度,削弱特征融合效果。为此,提出新型模态协作低秩分解框架(MC-LRD),从各模态中分解出不同领域偏移水平的模态独有与共享特征,更利于领域对齐。MC-LRD包含每模态多个分解器和多模态分解路由(MDR),分解器参数在不同模态间逐步共享;MDR选择性激活分解器以生成特征。为保证高效分解,对分解器与子路由分别施加正交去相关约束,增强多样性。此外,提出跨域激活一致性损失,确保同类别源/目标样本对分解器激活偏好一致,促进领域对齐。在三个公开基准上的实验表明,模型显著优于现有方法。
原文摘要 · Abstract (English)
In this paper, we study the challenging task of Few-Shot Video Domain Adaptation (FSVDA). The multimodal nature of videos introduces unique challenges, necessitating the simultaneous consideration of both domain alignment and modality collaboration in a few-shot scenario, which is ignored in previous literature. We observe that, under the influence of domain shift, the generalization performance on the target domain of each individual modality, as well as that of fused multimodal features, is constrained. Because each modality is comprised of coupled features with multiple components that exhibit different domain shifts. This variability increases the complexity of domain adaptation, thereby reducing the effectiveness of multimodal feature integration. To address these challenges, we introduce a novel framework of Modality-Collaborative LowRank Decomposers (MC-LRD) to decompose modality-unique and modality-shared features with different domain shift levels from each modality that are more friendly for domain alignment. The MC-LRD comprises multiple decomposers for each modality and Multimodal Decomposition Routers (MDR). Each decomposer has progressively shared parameters across different modalities. The MDR is leveraged to selectively activate the decomposers to produce modality-unique and modality-shared features. To ensure efficient decomposition, we apply orthogonal decorrelation constraints separately to decomposers and subrouters, enhancing their diversity. Furthermore, we propose a cross-domain activation consistency loss to guarantee that target and source samples of the same category exhibit consistent activation preferences of the decomposers, thereby facilitating domain alignment. Extensive experimental results on three public benchmarks demonstrate that our model achieves significant improvements over existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。