简单方法显著提升无监督视频域适应效果
Return of Frustratingly Easy Unsupervised Video Domain Adaptation

- 分空间与时间差异分别处理,用简洁架构实现
- 跨域动作识别任务上性能远超现有方法
- 适合做视频域适应的快速部署与基准测试
无监督视频域适应(UVDA)是一个实用但研究不足的问题。本文提出一种极简的UVDA方法MetaTrans,其学习目标仅包含两个基础损失项。尽管目标简单,但通过巧妙的模型设计,实现了对跨域视频空间与时间差异的分离处理。利用时间-静态减法模块,有效消除空间与时间偏差。大量实证评估表明,尤其在多种跨域动作识别任务中,相比当前最优的UVDA基线,性能有显著绝对提升和更优的相对增益。
原文摘要 · Abstract (English)
Unsupervised video domain adaptation (UVDA) is a practical but under-explored problem. In this paper, we propose a frustratingly easy UVDA method, called MetaTrans. Specifically, MetaTrans adopts a concise learning objective that contains only two fundamental loss terms. Despite the simplicity of the learning objective, MetaTrans embodies an advanced UVDA idea, that is, handling the spatial and temporal divergence of cross-domain videos separately, through a subtle model architecture design. By implementing a temporal-static subtraction module, MetaTrans effectively removes spatial and temporal divergence. Extensive empirical evaluations, particularly on various cross-domain action recognition tasks, show substantial absolute adaptation performance enhancement and significantly superior relative performance gain compared with state-of-the-art UVDA baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。