让视频模型专注动作而非物体,提升指令类视频理解能力
InstrAct: Towards Action-Centric Understanding in Instructional Videos
- 用动作中心的难样本过滤和对比学习,分离动作与物体干扰
- 引入动态时间对齐和掩码动作建模,强化时序结构与跨模态对齐
- 在指令视频任务中显著优于现有模型,适合需要精细动作理解场景
理解指令类视频需识别细粒度动作并建模其时序关系,当前视频基础模型(VFMs)面临挑战,根源在于网络监督噪声及普遍存在的“静态偏差”——模型依赖物体而非运动线索。为此,我们提出InstrAction预训练框架,用于构建指令视频的动作中心表征。首先设计数据驱动策略,过滤噪声字幕并生成动作中心的困难负样本,以在对比学习中解耦动作与物体。在视觉特征层面,引入动作感知器(Action Perceiver),从冗余视频编码中提取与运动相关标记。除对比学习外,还提出两项辅助目标:动态时间对齐(DTW-Align)以建模序列时序结构,掩码动作建模(MAM)以增强跨模态对齐。最后,构建InstrAct评测基准,验证方法在语义推理、流程逻辑和细粒度检索任务上持续优于现有最优的VFMs。
原文摘要 · Abstract (English)
Understanding instructional videos requires recognizing fine-grained actions and modeling their temporal relations, which remains challenging for current Video Foundation Models (VFMs). This difficulty stems from noisy web supervision and a pervasive "static bias", where models rely on objects rather than motion cues. To address this, we propose InstrAction, a pretraining framework for instructional videos' action-centric representations. We first introduce a data-driven strategy, which filters noisy captions and generates action-centric hard negatives to disentangle actions from objects during contrastive learning. At the visual feature level, an Action Perceiver extracts motion-relevant tokens from redundant video encodings. Beyond contrastive learning, we introduce two auxiliary objectives: Dynamic Time Warping alignment (DTW-Align) for modeling sequential temporal structure, and Masked Action Modeling (MAM) for strengthening cross-modal grounding. Finally, we introduce the InstrAct Bench to evaluate action-centric understanding, where our method consistently outperforms state-of-the-art VFMs on semantic reasoning, procedural logic, and fine-grained retrieval tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。