arXiv:2505.19495cs.CV2025-05IJCAI被引 1

用文本生成视频数据,解决动作识别中样本不足的问题

The Role of Video Generation in Enhancing Data-Limited Action Understanding

  • 用文本到视频扩散模型自动生成带标注数据
  • 在5个任务上实现零样本动作识别新纪录
  • 通过增强信息和不确定度平滑提升生成数据质量

真实场景中的视频动作理解常受数据稀缺制约。本文提出一种新方法,利用文本到视频的扩散变换器生成无限量无须人工干预的带标注训练数据。定量与定性分析表明,真实样本通常比生成样本包含更丰富的信息。为此,我们提出信息增强策略,从环境和角色两方面提升生成样本的信息量;同时观察到部分低质生成样本可能干扰模型训练,因此设计基于不确定度的标签平滑策略,降低其影响。该方法在四个数据集上的五个任务中均验证有效,实现零样本动作识别的最新性能。

原文摘要 · Abstract (English)

Video action understanding tasks in real-world scenarios always suffer data limitations. In this paper, we address the data-limited action understanding problem by bridging data scarcity. We propose a novel method that employs a text-to-video diffusion transformer to generate annotated data for model training. This paradigm enables the generation of realistic annotated data on an infinite scale without human intervention. We proposed the information enhancement strategy and the uncertainty-based label smoothing tailored to generate sample training. Through quantitative and qualitative analysis, we observed that real samples generally contain a richer level of information than generated samples. Based on this observation, the information enhancement strategy is proposed to enhance the informative content of the generated samples from two aspects: the environments and the characters. Furthermore, we observed that some low-quality generated samples might negatively affect model training. To address this, we devised the uncertainty-based label smoothing strategy to increase the smoothing of these samples, thus reducing their impact. We demonstrate the effectiveness of the proposed method on four datasets across five tasks and achieve state-of-the-art performance for zero-shot action recognition.

视频生成动作识别数据增强扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。