arXiv:2502.20811cs.CVcs.CL2025-02ACL被引 2

用高质量描述提升多模态模型对人类动作的理解与生成能力

HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models

  • 设计两阶段标注流程,用人体特征和时间顺序精准描述动作
  • 构建12.6万条训练数据,412条人工标注测试集,涵盖问答对2000组
  • 显著提升4个基准上的动作理解与文本生成效果,适合视频智能研究者

近期多模态大模型在视频理解上取得进展,但在涉及人类动作的视频任务中仍受限于高质量数据不足。为此,我们提出两阶段数据标注流程:首先从网络收集清晰的人类动作视频,其次采用标准化标题格式,利用人体属性区分个体,并按时间顺序详细描述动作与交互。基于此流程,我们构建了两个数据集:HAICTrain包含12.6万条由Gemini-Pro生成并经验证的视频-标题对,用于训练;HAICBench包含412条人工标注的视频-标题对及2000个问答对,用于全面评估人类动作理解能力。实验表明,使用HAICTrain训练不仅能显著提升4个基准上的表现,还能改善文本到视频生成效果。两个数据集已公开于https://huggingface.co/datasets/KuaishouHAIC/HAIC。

原文摘要 · Abstract (English)

Recent Multi-modal Large Language Models (MLLMs) have made great progress in video understanding. However, their performance on videos involving human actions is still limited by the lack of high-quality data. To address this, we introduce a two-stage data annotation pipeline. First, we design strategies to accumulate videos featuring clear human actions from the Internet. Second, videos are annotated in a standardized caption format that uses human attributes to distinguish individuals and chronologically details their actions and interactions. Through this pipeline, we curate two datasets, namely HAICTrain and HAICBench. \textbf{HAICTrain} comprises 126K video-caption pairs generated by Gemini-Pro and verified for training purposes. Meanwhile, \textbf{HAICBench} includes 412 manually annotated video-caption pairs and 2,000 QA pairs, for a comprehensive evaluation of human action understanding. Experimental results demonstrate that training with HAICTrain not only significantly enhances human understanding abilities across 4 benchmarks, but can also improve text-to-video generation results. Both the HAICTrain and HAICBench are released at https://huggingface.co/datasets/KuaishouHAIC/HAIC.

动作理解多模态数据集生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。