提出两种多视角技能评估方法,两阶段模型表现更优。
CuriosAI Submission to the EgoExo4D Proficiency Estimation Challenge 2025
- 用Sapiens-2B做多任务学习,同时预测技能水平和场景类型。
- 两阶段方法结合零样本场景识别与视图专用分类器,准确率达47.8%。
- 适合关注多视角技能评估与场景感知的视觉算法研究者。
本文介绍CuriosAI团队在CVPR 2025举办的EgoExo4D技能评估挑战赛中的参赛方案。提出两种多视角技能评估方法:(1) 基于Sapiens-2B的多任务学习框架,联合预测技能水平与场景标签,准确率为43.6%;(2) 两阶段流水线,先进行零样本场景识别,再使用视图专用的VideoMAE分类器进行评估,准确率达到47.8%。结果表明,基于场景条件建模的两阶段方法性能更优,验证了场景信息对技能评估的关键作用。
原文摘要 · Abstract (English)
This report presents the CuriosAI team's submission to the EgoExo4D Proficiency Estimation Challenge at CVPR 2025. We propose two methods for multi-view skill assessment: (1) a multi-task learning framework using Sapiens-2B that jointly predicts proficiency and scenario labels (43.6 % accuracy), and (2) a two-stage pipeline combining zero-shot scenario recognition with view-specific VideoMAE classifiers (47.8 % accuracy). The superior performance of the two-stage approach demonstrates the effectiveness of scenario-conditioned modeling for proficiency estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。