arXiv:2409.18478cs.CV2024-09中稿 · CVIU

统一视频时间理解任务,用离散序列表示实现多任务通用建模。

Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks

  • 将时序动作检测等任务输出转为离散令牌序列,构建统一框架。
  • 在三个任务上表现合理,优于单一任务训练模型。
  • 通用模型在新数据集上泛化能力更强,适合多任务研究者。

随着视频理解的发展,片段级时序视频分析任务日益增多,包括时序动作检测(TAD)、时序动作分割(TAS)和通用事件边界检测(GEBD)。尽管特定任务的模型表现优异,但缺乏能同时处理多项任务的统一框架,这成为下一代AI的重要方向。为此,本文提出统一框架Temporal2Seq,将这些任务的输出形式化为离散令牌序列。通过该统一表示,Temporal2Seq可在单一架构内训练通用模型。由于缺乏多任务学习基准,我们从TAD、TAS和GEBD任务中借用数据构建综合共训练数据集。在三个任务的测试集上评估,结果表明Temporal2Seq在各类任务中表现良好,并优于该框架下的单任务训练模型。此外,通用模型在不同任务的新数据集上也展现出更优的泛化性能,优于专用模型。

原文摘要 · Abstract (English)

With the development of video understanding, there is a proliferation of tasks for clip-level temporal video analysis, including temporal action detection (TAD), temporal action segmentation (TAS), and generic event boundary detection (GEBD). While task-specific video understanding models have exhibited outstanding performance in each task, there remains a dearth of a unified framework capable of simultaneously addressing multiple tasks, which is a promising direction for the next generation of AI. To this end, in this paper, we propose a single unified framework, coined as Temporal2Seq, to formulate the output of these temporal video understanding tasks as a sequence of discrete tokens. With this unified token representation, Temporal2Seq can train a generalist model within a single architecture on different video understanding tasks. In the absence of multi-task learning (MTL) benchmarks, we compile a comprehensive co-training dataset by borrowing the datasets from TAD, TAS, and GEBD tasks. We evaluate our Temporal2Seq generalist model on the corresponding test sets of three tasks, demonstrating that Temporal2Seq can produce reasonable results on various tasks and achieve advantages compared with single-task training on this framework. We also investigate the generalization performance of our generalist model on new datasets from different tasks, which yields superior performance to the specific model.

视频理解多任务学习时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。