arXiv:2605.20904cs.CV2026-05被引 2

基于JEPA的未来动作预测模型,登顶EgoVis 2026挑战赛

JFAA: Technical Report for the EPIC-KITCHENS-100 Action Anticipation Challenge at EgoVis 2026

论文配图:JFAA: Technical Report for the EPIC-KITCHENS-100 Action Anticipation Challenge at EgoVis 2026
图 1 · 摘自论文原文
  • 用冻结编码器和预测器提取上下文与近未来特征
  • 轻量注意力探针分别预测动词、名词和动作类别
  • 场域感知集成提升鲁棒性,适合视觉动作预测研究者

我们提出JFAA,一种基于JEPA的未来动作预测方法,用于EPIC-KITCHENS-100(EK-100)动作预测任务。受V-JEPA 2.1表示学习与未来预测能力启发,JFAA采用冻结编码器和预测器提取观测上下文特征与近未来潜在标记。随后训练一个轻量级注意力探针,通过独立任务查询预测动词、名词和动作的分类得分。为增强鲁棒性,进一步构建基于选定轮次预测结果的场域感知集成策略,使每个输出字段可利用最可靠的候选结果。在官方挑战服务器上的实验表明,JFAA在EgoVis 2026 EK-100动作预测挑战中取得第一名。代码将发布于https://github.com/CorrineQiu/JFAA。

原文摘要 · Abstract (English)

We propose JFAA, a JEPA-based Future Action Anticipation method for the EPIC-KITCHENS-100 (EK-100) Action Anticipation task. Inspired by the representation learning and future prediction ability of V-JEPA 2.1, JFAA uses a frozen encoder and predictor to extract observed context features and near-future latent tokens. A lightweight attentive probe is then trained to predict verb, noun, and action logits with separate task queries. To improve robustness, we further build a field-aware ensemble over selected epoch-level predictions, allowing each output field to benefit from its most reliable candidates. Experimental results on the official challenge server show that JFAA achieves first place in the EgoVis 2026 EK-100 Action Anticipation Challenge. Our code will be released at https://github.com/CorrineQiu/JFAA.

动作预测自监督学习视频理解JEPA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。