arXiv:2504.06580cs.CVcs.AI2025-04中稿 · ICLR

发现动作识别模型依赖固定顺序,提出两种方法提升泛化能力

Exploring Ordinal Bias in Action Recognition for Instructional Videos

  • 通过掩码高频共现动作帧和打乱动作顺序来削弱顺序依赖
  • 模型在非标准动作序列上性能显著下降,暴露顺序偏差问题
  • 适合关注视频理解泛化性与评估方式改进的研究者

动作识别模型在理解教学视频方面已取得良好效果,但往往依赖数据集中固定的动作序列,而非真正理解视频内容,这种现象我们称为顺序偏差。为解决该问题,本文提出两种有效的视频操作方法:动作掩码(Action Masking),即遮蔽频繁共现动作的帧;序列打乱(Sequence Shuffling),即随机化动作片段的顺序。通过全面实验,我们发现当前模型在面对非标准动作序列时性能大幅下降,凸显其对顺序偏差的敏感性。研究强调需重新思考评估策略,并开发能超越固定动作模式、适应多样化教学视频的模型。

原文摘要 · Abstract (English)

Action recognition models have achieved promising results in understanding instructional videos. However, they often rely on dominant, dataset-specific action sequences rather than true video comprehension, a problem that we define as ordinal bias. To address this issue, we propose two effective video manipulation methods: Action Masking, which masks frames of frequently co-occurring actions, and Sequence Shuffling, which randomizes the order of action segments. Through comprehensive experiments, we demonstrate that current models exhibit significant performance drops when confronted with nonstandard action sequences, underscoring their vulnerability to ordinal bias. Our findings emphasize the importance of rethinking evaluation strategies and developing models capable of generalizing beyond fixed action patterns in diverse instructional videos.

动作识别顺序偏差视频理解泛化性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。