构建细粒度动作数据集,提升模型对人类动作的精准理解能力
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
- 构建包含数千视频的细粒度动作标注数据集,精确标记每处肢体运动
- 现有大模型在细粒度任务上表现不佳,主要因缺乏精细标注数据
- 提出自动生成的代理任务,降低人工标注依赖,显著提升模型感知能力
视频中人类动作与姿态的细粒度理解对以人为本的AI应用至关重要。本文提出ActionArt,一个面向人类中心多模态理解的细粒度视频-文本数据集。该数据集包含数千个涵盖广泛人类动作、人机交互及多样场景的视频,每个视频均配有详细标注,精确记录每一肢体的运动。我们设计了八个子任务,用于评估现有大型多模态模型在不同维度上的细粒度理解能力。实验表明,尽管当前大模型在多数任务上表现良好,但在细粒度理解方面仍存在明显不足。我们认为这一局限源于高质量精细标注数据稀缺,而此类标注成本高且难以规模化。由于人工标注成本高昂且难扩展,我们提出了由现有多模态大模型(MLLMs)自动生成数据驱动的代理任务,以增强模型在时空维度上的感知能力。实验结果表明,所提代理任务显著缩小了与使用人工标注数据时性能之间的差距。
原文摘要 · Abstract (English)
Fine-grained understanding of human actions and poses in videos is essential for human-centric AI applications. In this work, we introduce ActionArt, a fine-grained video-caption dataset designed to advance research in human-centric multimodal understanding. Our dataset comprises thousands of videos capturing a broad spectrum of human actions, human-object interactions, and diverse scenarios, each accompanied by detailed annotations that meticulously label every limb movement. We develop eight sub-tasks to evaluate the fine-grained understanding capabilities of existing large multimodal models across different dimensions. Experimental results indicate that, while current large multimodal models perform commendably on various tasks, they often fall short in achieving fine-grained understanding. We attribute this limitation to the scarcity of meticulously annotated data, which is both costly and difficult to scale manually. Since manual annotations are costly and hard to scale, we propose proxy tasks to enhance the model perception ability in both spatial and temporal dimensions. These proxy tasks are carefully crafted to be driven by data automatically generated from existing MLLMs, thereby reducing the reliance on costly manual labels. Experimental results show that the proposed proxy tasks significantly narrow the gap toward the performance achieved with manually annotated fine-grained data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。