用多视角视频生成专业级动作评价反馈,更智能且省资源。
ProfVLM: A lightweight video-language model for multi-view proficiency estimation
- 将动作评估转为生成式建模,结合多视角视频输出评分与自然语言反馈。
- 在EgoExo4D数据集上超越现有方法,参数量少20倍,训练快60%。
- 适合需要可解释性动作评估的体育教学、康复训练等场景。
现有方法将动作质量评估和技能水平估计视为分类任务,通常仅输出离散标签或分数,未显式建模评估背后的推理过程。本文提出一种生成式视觉-语言建模框架ProfVLM,通过联合预测熟练度等级并生成类专家的自然语言反馈,实现可解释的多视角动作评估。ProfVLM利用动态注意力门控投影器(AttentiveGatedProjector),将冻结的TimeSformer主干网络提取的多视角第一人称与第三人称特征融合并投影至微调后的语言模型,以生成反馈。该模型在包含专家评语的EgoExo4D数据集上训练,性能超越现有最先进方法,同时参数量最多减少20倍,训练时间缩短达60%。结果表明,生成式视觉-语言建模为可解释的动作质量评估提供了一种高效的新范式。
原文摘要 · Abstract (English)
Most existing approaches formulate action quality assessment and skill proficiency estimation as discriminative prediction tasks, typically producing discrete labels or scores without explicitly modeling the reasoning process underlying the assessment. We instead reformulate the problem as generative vision-language modeling, introducing ProfVLM, a parameter-efficient vision-language model that jointly predicts proficiency levels and generates expert-like natural language feedback from multi-view videos. ProfVLM leverages conditional language generation to provide actionable insights along with quantitative evaluation scores. Central to our method is an AttentiveGatedProjector that dynamically fuses and projects multi-view egocentric and exocentric features from a frozen TimeSformer backbone into a language model fine-tuned for feedback generation. Trained on EgoExo4D with expert commentaries, ProfVLM surpasses state-of-the-art methods while using up to 20x fewer parameters and reducing training time by up to 60% compared to existing classification-based methods. By providing natural language critiques aligned with performance levels, this work shows that generative vision-language modeling offers a powerful and efficient paradigm shift for interpretable action quality assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。