用少参数实现多视角动作熟练度评估,还能生成专家反馈。
Parameter-Efficient Multi-View Proficiency Estimation: From Discriminative Classification to Generative Feedback

- 通过选择性融合多视角信息,减少参数量
- 在Ego-Exo4D上达最优精度,参数减少20倍
- 从分类转向生成式反馈,适合教练与康复场景
评估动作熟练度而非动作类别,对指导、康复和人才选拔至关重要。该任务挑战在于熟练度隐含在时间、平衡、身体力学等细微差异中,常分布于多个视角和短暂事件。本文在Ego-Exo4D数据集上提出三项贡献:SkillFormer采用参数高效判别架构实现选择性多视角融合;PATS通过保留关键动作的密集局部片段改善时间采样;ProfVLM将熟练度评估重构为条件语言生成,借助门控跨视角投影器与紧凑语言骨干,输出熟练度标签及专家风格反馈。三者结合在Ego-Exo4D上达到当前最优性能,相比视频变换器基线,参数量减少最多20倍,训练轮次减少最多3倍,且从封闭集分类转向可解释反馈生成。结果表明,高效、多视角系统正向选择性融合、熟练度感知采样与可行动生成反馈演进。
原文摘要 · Abstract (English)
Estimating how well a person performs an action, rather than which action is performed, is central to coaching, rehabilitation, and talent identification. This task is challenging because proficiency is encoded in subtle differences in timing, balance, body mechanics, and execution, often distributed across multiple views and short temporal events. We discuss three recent contributions to multi-view proficiency estimation on Ego-Exo4D. SkillFormer introduces a parameter-efficient discriminative architecture for selective multi-view fusion; PATS improves temporal sampling by preserving locally dense excerpts of fundamental movements; and ProfVLM reformulates proficiency estimation as conditional language generation, producing both a proficiency label and expert-style feedback through a gated cross-view projector and a compact language backbone. Together, these methods achieve state-of-the-art accuracy on Ego-Exo4D with up to 20x fewer trainable parameters and up to 3x fewer training epochs than video-transformer baselines, while moving from closed-set classification toward interpretable feedback generation. These results highlight a shift toward efficient, multi-view systems that combine selective fusion, proficiency-aware sampling, and actionable generative feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。