arXiv:2409.09306cs.CV2024-09中稿 · the International …被引 3

用人体关键点增强视觉指令数据,提升模型对人姿动作的理解能力

Keypoint-Integrated Instruction-Following Data Generation for Enhanced Human Pose and Action Understanding in Multimodal Models

  • 将人体关键点与图像描述、框选信息结合生成指令数据
  • 在20万样本数据上微调后,性能提升21.18%
  • 适合需要精准理解人体姿态与动作的研究者

当前视觉语言多模态模型在通用视觉理解任务中表现良好,但在涉及人体姿态与动作的复杂任务中表现不足,主要因缺乏专门的视觉语言指令跟随数据。本文提出一种方法,将人体关键点与传统视觉特征(如字幕、边界框)融合,生成更精确的人类中心场景理解数据。构建了一个包含200,328个样本的数据集,用于微调模型以应对对话、详细描述和复杂推理三类任务。为此建立基准测试集Human Pose and Action Understanding Benchmark (HPAUB),使用该数据集对LLaVA-1.5-7B模型进行微调,并在基准上评估,相比原模型整体性能提升21.18%。结果表明,关键点融合数据能有效提升多模态模型对人类姿态与动作的理解能力。代码已开源:https://github.com/Ody-trek/Keypoint-Instruction-Tuning。

原文摘要 · Abstract (English)

Current vision-language multimodal models are well-adapted for general visual understanding tasks. However, they perform inadequately when handling complex visual tasks related to human poses and actions due to the lack of specialized vision-language instruction-following data. We introduce a method for generating such data by integrating human keypoints with traditional visual features such as captions and bounding boxes, enabling more precise understanding of human-centric scenes. Our approach constructs a dataset comprising 200,328 samples tailored to fine-tune models for human-centric tasks, focusing on three areas: conversation, detailed description, and complex reasoning. We establish a benchmark called Human Pose and Action Understanding Benchmark (HPAUB) to assess model performance on human pose and action understanding. We fine-tune the LLaVA-1.5-7B model using this dataset and evaluate it on the benchmark, achieving significant improvements. Experimental results show an overall improvement of 21.18% compared to the original LLaVA-1.5-7B model. These findings highlight the effectiveness of keypoint-integrated data in enhancing multimodal models. Code is available at https://github.com/Ody-trek/Keypoint-Instruction-Tuning.

人体姿态指令微调多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。