arXiv:2505.08665cs.CV2025-05中稿 · the 2025 18th Inte…被引 7

用多视角视频统一评估人类技能水平,效率更高精度更强。

SkillFormer: Unified Multi-View Video Understanding for Proficiency Estimation

  • 通过跨视角融合模块整合第一/第三人称视频特征
  • 在EgoExo4D数据集上准确率领先,参数少4.5倍、训练快3.75倍
  • 适合需要精细动作评估的体育与康复场景

评估复杂活动中的人类技能水平是一项具有挑战性的任务,应用广泛于体育、康复和训练领域。本文提出SkillFormer,一种基于TimeSformer架构的参数高效统一多视角技能评估模型。该模型引入跨视角融合模块,利用多头交叉注意力、可学习门控和自适应校准机制融合视图特异性特征,并采用低秩适配技术仅微调少量参数,显著降低训练成本。在EgoExo4D数据集上的实验表明,SkillFormer在多视角设置下达到当前最优准确率,同时参数量减少4.5倍,训练轮次减少3.75倍。其在多项结构化任务中表现优异,验证了多视角信息融合对细粒度技能评估的价值。

原文摘要 · Abstract (English)

Assessing human skill levels in complex activities is a challenging problem with applications in sports, rehabilitation, and training. In this work, we present SkillFormer, a parameter-efficient architecture for unified multi-view proficiency estimation from egocentric and exocentric videos. Building on the TimeSformer backbone, SkillFormer introduces a CrossViewFusion module that fuses view-specific features using multi-head cross-attention, learnable gating, and adaptive self-calibration. We leverage Low-Rank Adaptation to fine-tune only a small subset of parameters, significantly reducing training costs. In fact, when evaluated on the EgoExo4D dataset, SkillFormer achieves state-of-the-art accuracy in multi-view settings while demonstrating remarkable computational efficiency, using 4.5x fewer parameters and requiring 3.75x fewer training epochs than prior baselines. It excels in multiple structured tasks, confirming the value of multi-view integration for fine-grained skill assessment. Project page at https://edowhite.github.io/SkillFormer

视频理解技能评估多视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。