通过多尺度时间对齐,提升跨域视频动作识别效果。
Multi-Scale Temporal Domain Alignment for Federated Video Domain Adaptation

- 在多个时间粒度上训练变换器编码器,实现跨客户端知识融合。
- 在Epic-Kitchens-55和Daily-DA上达到28.47%的性能提升。
- 适合隐私保护下的分布式视频分析场景,尤其关注时间建模。
联邦视频域适应(FVDA)可在保护隐私的前提下协同学习分布异构的视频数据,但因时间信息对齐困难而研究较少。本文提出多尺度时间域对齐(METAL)框架,利用多分辨率时间信息提升跨域视频动作识别性能,仅通过模型参数传输实现协作。METAL在源端客户端分别训练各尺度的Transformer编码器,目标服务器独立进行多尺度知识投票生成鲁棒伪标签。引入新型$L_2$方差惩罚项,在尺度知识蒸馏过程中强制跨尺度一致性,避免单一主导尺度。特征通过后期融合聚合,融合头基于置信度加权的尺度预测进行知识蒸馏,有效挖掘互补的时间信息以支持最终预测。在Epic-Kitchens-55与Daily-DA数据集上的实验表明,该方法达到当前最优性能,相较现有FDA方法最高提升28.47%。消融实验验证多尺度蒸馏与尺度协调对有效时间知识迁移至关重要。
原文摘要 · Abstract (English)
Federated Video Domain Adaptation (FVDA) enables collaborative learning across distributed and non-IID video datasets while preserving privacy, but is under-explored due to challenges in aligning temporal information. We propose Multi-scalE Temporal domAin aLignment (METAL), a novel framework that leverages temporal information at multiple resolutions to improve cross-domain video action recognition with only model parameter transfers. METAL trains per-scale transformer encoders on source-clients, then performs independent knowledge voting at each temporal scale to generate robust pseudo-labels on the target-server. A novel $L_2$ variance penalty enforces cross-scale consistency during scale-based knowledge distillation, preventing a singular dominant scale. The late fusion aggregates features across different scales, where the fusion head is trained via knowledge distillation using confidence-weighted aggregation of scale-wise predictions, enabling the model to effectively exploit complementary temporal information for final predictions. Experiments on Epic-Kitchens-55 and Daily-DA demonstrate state-of-the-art performances, with gains up to 28.47% over current FDA methods. Ablation studies prove that multi-scale distillation and scale coordination are critical for effective temporal knowledge transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。