用知识蒸馏提升少样本精准事件定位,视频与骨骼信息协同优化
From Skeletons to Pixels: Few-Shot Precise Event Spotting via Representation and Prediction Distillation

- 通过自适应加权和渐进伪标签实现多模态知识蒸馏
- 在网球和花样滑冰数据集上,小样本下准确率显著超越基线
- 适合运动分析、视频理解等少标注场景的精准事件识别
精确事件定位(PES)在网球等快节奏体育赛事中至关重要,因细粒度事件发生在极短时间内,帧级精确定位面临运动模糊、动作细微差异及标注数据有限等挑战。本文研究两种互补的知识蒸馏策略:预测层的自适应权重蒸馏(AWD),在无标注数据上动态调整教师模型监督;以及表示层的渐进多模态蒸馏(AMD-FED),通过渐进伪标签将鲁棒的骨骼知识迁移至视觉模态。二者均采用多模态蒸馏,在少样本k-clip设置下于F3Set-Tennis(sub)数据集上持续优于单模态基线与现有PES方法。进一步在花样滑冰数据集验证表明,AMD-FED在少样本下仍具强鲁棒性。结果表明,尤其表示层知识迁移在少样本精准事件定位中效果显著。
原文摘要 · Abstract (English)
Precise Event Spotting (PES) is essential in fast-paced sports such as tennis, where fine-grained events occur within very short temporal windows. Accurate frame-level localization is challenging because of motion blur, subtle action differences, and limited annotated data. We study two complementary distillation strategies for few-shot PES: Adaptive Weight Distillation (AWD), a prediction-level method that adaptively weights teacher supervision on unlabeled data, and Annealed Multimodal Distillation for Few-Shot Event Detection (AMD-FED), a representation-level framework that transfers robust skeleton knowledge into visual modalities through annealed pseudo-labeling. Both methods use multimodal distillation to improve generalization under limited supervision. We evaluate them on F3Set-Tennis(sub) under few-shot k-clip settings, where they consistently outperform single-modality baselines and prior PES approaches. After observing the stronger performance of representation-level distillation on tennis, we further validate AMD-FED on a second sports dataset, Figure Skating, where it also shows robust performance in the k-clip scenario. These results highlight the effectiveness of multimodal distillation, especially representation-level transfer, for few-shot precise event spotting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。