arXiv:2604.09955cs.CV2026-04中稿 · IEEE/CVF Conferenc…

通过动态筛选运动信息丰富的视频片段,提升无监督域适应的准确率与效率。

Learnable Motion-Focused Tokenization for Effective and Efficient Video Unsupervised Domain Adaptation

论文配图:Learnable Motion-Focused Tokenization for Effective and Efficient Video Unsupervised Domain Adaptation
图 1 · 摘自论文原文
  • 自学习提取运动关键区域,过滤静态背景冗余信息
  • 在21组跨域设置下超越现有方法,计算量大幅降低
  • 适合需高效部署的视频动作识别场景

视频无监督域适应(VUDA)在动作识别中面临挑战,需将带标签源域模型迁移到无标签目标域。尽管近期有进展,现有方法仍难以达到全监督性能,主要原因在于静态、低信息量背景加剧了域偏移。同时,以往方法普遍忽视计算效率,限制实际应用。为此,本文提出可学习的运动聚焦分块(LMFT)用于VUDA:将视频帧划分为块令牌,自学习剔除低运动冗余块(主要对应背景),保留富含运动信息的动作相关块进行适应。在三个标准VUDA基准上21组域迁移设置的大量实验表明,采用LMFT的框架实现最先进性能,同时显著降低计算开销。因此,LMFT实现了既有效又高效的VUDA。

原文摘要 · Abstract (English)

Video Unsupervised Domain Adaptation (VUDA) poses a significant challenge in action recognition, requiring the adaptation of a model from a labeled source domain to an unlabeled target domain. Despite recent advances, existing VUDA methods often fall short of fully supervised performance, a key reason being the prevalence of static and uninformative backgrounds that exacerbate domain shifts. Additionally, prior approaches largely overlook computational efficiency, limiting real-world adoption. To address these issues, we propose Learnable Motion-Focused Tokenization (LMFT) for VUDA. LMFT tokenizes video frames into patch tokens and learns to discard low-motion, redundant tokens, primarily corresponding to background regions, while retaining motion-rich, action-relevant tokens for adaptation. Extensive experiments on three standard VUDA benchmarks across 21 domain adaptation settings show that our VUDA framework with LMFT achieves state-of-the-art performance while significantly reducing computational overhead. LMFT thus enables VUDA that is both effective and computationally efficient.

视频域适应运动聚焦计算效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。