arXiv:2502.17352cs.CV2025-02被引 1

利用任务层级结构提升视频预训练效率,更好识别步骤与任务。

Leveraging Procedural Knowledge and Task Hierarchies for Efficient Instructional Video Pre-training

  • 引入任务层级和流程步骤作为先验知识,指导模型预训练。
  • 在有限数据和算力下,任务识别与步骤预测性能超越现有基线。
  • 适合资源受限场景下的视频推荐系统开发人员使用。

教学视频为学习新任务(如烹饪食谱或组装家具)提供了便捷的媒介。观众希望找到既符合总体任务目标,又包含所需具体步骤的视频。为此,教学视频模型需能推断输入视频中的任务和步骤。当计算资源或训练数据有限时,高效且可泛化的建模尤为关键。为此,我们显式挖掘教学视频中的任务层级及对应流程步骤,并利用这些先验知识对模型 $ exttt{Pivot}$ 进行步骤与任务预测的预训练。预训练过程中,采用视频增强与早停策略,以最优选择下游任务适用的模型。我们在两个下游数据集上测试了该模型在任务识别、步骤识别和步骤预测任务上的表现。在预训练数据与算力受限条件下,$ exttt{Pivot}$ 在各项任务中均优于先前基线。因此,利用先验任务与步骤结构,能够高效训练 $ exttt{Pivot}$ 用于教学视频推荐。

原文摘要 · Abstract (English)

Instructional videos provide a convenient modality to learn new tasks (ex. cooking a recipe, or assembling furniture). A viewer will want to find a corresponding video that reflects both the overall task they are interested in as well as contains the relevant steps they need to carry out the task. To perform this, an instructional video model should be capable of inferring both the tasks and the steps that occur in an input video. Doing this efficiently and in a generalizable fashion is key when compute or relevant video topics used to train this model are limited. To address these requirements we explicitly mine task hierarchies and the procedural steps associated with instructional videos. We use this prior knowledge to pre-train our model, $\texttt{Pivot}$, for step and task prediction. During pre-training, we also provide video augmentation and early stopping strategies to optimally identify which model to use for downstream tasks. We test this pre-trained model on task recognition, step recognition, and step prediction tasks on two downstream datasets. When pre-training data and compute are limited, we outperform previous baselines along these tasks. Therefore, leveraging prior task and step structures enables efficient training of $\texttt{Pivot}$ for instructional video recommendation.

视频理解任务识别预训练教学视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。