统一多视角技能评估框架,用动态路由提升跨场景识别准确率。
SkillMoV: Mixture-of-View Routing with Prototype-Conditioned Gating for Unified Multi-View Proficiency Estimation

- 用多视角软路由和原型条件门控,自适应融合不同摄像头的特征。
- 在六个领域上达到50.17%准确率,比现有方法高3.57个百分点。
- 参数高效训练仅需23.32%参数量,适合实际部署与快速适配。
从视频中估计人类技能水平是自动化技能评估的关键挑战,广泛应用于体育训练、音乐教学、外科培训和职场学习。现有方法通常局限于单一场景或依赖共享多视角聚合,难以适应异构摄像机视角和活动领域。本文提出SkillMoV,一个统一且参数高效的多场景技能评估框架,可处理同步多视角视频。核心为多视角投影器(MoVP),采用混合专家范式适配相机特定特征,包含四阶段:(i) 十二个专家MLP的多视角软路由,无相机身份监督下学习视角依赖的专家偏好;(ii) 跨视角注意力对齐同步摄像头;(iii) 可学习的原型锚定,以类别级参考向量条件化表示;(iv) 原型条件门控投影生成最终技能嵌入。在EgoExo4D数据集上评估,覆盖六个技能领域和三种视角配置(Ego、Exos、Ego+Exos)。SkillMoV在Exos设置中实现50.17%总体准确率,单模型联合训练超越对比方法3.57个百分点;在Ego+Exos中达47.63%,接近最佳结果(48.20%)。消融实验验证各组件贡献:MoV路由 +6.61个百分点,跨视角注意力 +4.92,原型锚定 +4.07,随机视角丢弃 +3.90。通过LoRA适配,仅训练23.32%参数,性能开销极低。
原文摘要 · Abstract (English)
Estimating human proficiency from video is a key challenge for automated skill assessment, with applications in sports coaching, music pedagogy, surgical training, and workplace learning. Existing approaches often focus on individual scenarios or rely on shared multi-view aggregation, limiting their ability to adapt to heterogeneous camera viewpoints and activity domains. We introduce SkillMoV, a unified, parameter-efficient framework for multi-scenario proficiency estimation from synchronized multi-view video. At its core, SkillMoV introduces a Mixture-of-View Projector (MoVP), which adapts the mixture-of-experts paradigm to camera-specific view features. MoVP is composed of four stages: (i) a Mixture-of-View soft router with twelve expert MLPs that learns view-dependent expert preferences without camera-identity supervision; (ii) cross-view attention to align synchronized cameras; (iii) learnable prototype anchoring to condition the representation on class-level reference vectors; and (iv) a prototype-conditioned gated projection that produces the final skill embedding. We evaluate SkillMoV on EgoExo4D across six skill domains and three separately trained view configurations: Ego, Exos, and Ego+Exos. SkillMoV reaches 50.17% overall accuracy in the Exos setting with a single model trained jointly across all scenarios, surpassing the strongest reported Exos result among the compared methods by 3.57 percentage points. In Ego+Exos, SkillMoV remains close to the best reported result in that setting (47.63% versus 48.20%). Ablations on the selected Exos configuration validate each component: MoV routing contributes +6.61 pp over attentive aggregation, cross-view attention +4.92 pp, prototype anchoring +4.07 pp, and stochastic view dropout +3.90 pp. Through LoRA adaptation, SkillMoV trains only 23.32% of its parameters and adds limited measured overhead relative to a LoRA-only baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。