arXiv:2511.12870cs.CV2025-11被引 2

提出跨视角知识蒸馏框架,让弱监督学生模型超越强监督教师。

View-aware Cross-modal Distillation for Multi-view Action Recognition

  • 设计跨模态适配器,利用不完整模态中的多模态关联
  • 通过视图一致性模块对齐部分可见动作的预测,提升鲁棒性
  • 在真实场景数据集上显著优于现有方法,适合资源受限场景

多传感器系统普及推动了多视角动作识别研究。现有方法在全重叠视角下表现良好,但部分重叠场景(动作仅在部分视角可见)仍待探索。现实系统常受限于输入模态和序列级标注,缺乏密集帧级标签。本文提出视图感知跨模态知识蒸馏(ViCoKD),将全监督多模态教师的知识迁移到模态与标注均受限的学生模型。ViCoKD采用带跨模态注意力的跨模态适配器,使学生在不完整模态下仍能利用多模态相关性。此外,提出视图感知一致性模块,解决视图间动作表达差异或部分可见问题:当动作在多个视角共现时,基于人体检测掩码和置信加权的JS散度对预测进行对齐。在真实世界数据集MultiSensor-Home上的实验表明,ViCoKD在多种骨干网络与环境下持续优于对比方法,在有限条件下甚至超越教师模型。

原文摘要 · Abstract (English)

The widespread use of multi-sensor systems has increased research in multi-view action recognition. While existing approaches in multi-view setups with fully overlapping sensors benefit from consistent view coverage, partially overlapping settings where actions are visible in only a subset of views remain underexplored. This challenge becomes more severe in real-world scenarios, as many systems provide only limited input modalities and rely on sequence-level annotations instead of dense frame-level labels. In this study, we propose View-aware Cross-modal Knowledge Distillation (ViCoKD), a framework that distills knowledge from a fully supervised multi-modal teacher to a modality- and annotation-limited student. ViCoKD employs a cross-modal adapter with cross-modal attention, allowing the student to exploit multi-modal correlations while operating with incomplete modalities. Moreover, we propose a View-aware Consistency module to address view misalignment, where the same action may appear differently or only partially across viewpoints. It enforces prediction alignment when the action is co-visible across views, guided by human-detection masks and confidence-weighted Jensen-Shannon divergence between their predicted class distributions. Experiments on the real-world MultiSensor-Home dataset show that ViCoKD consistently outperforms competitive distillation methods across multiple backbones and environments, delivering significant gains and surpassing the teacher model under limited conditions.

动作识别知识蒸馏多视角弱监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。