arXiv:2411.03034cs.AIcs.MM2024-11被引 22

专为人体场景设计的视觉语言模型,提升人相关任务表现

HumanVLM: Foundation for Human-Scene Vision-Language Model

  • 构建人体场景多模态数据集,强化人与环境的对齐
  • 在31.1万张图像上训练,显著提升人相关任务性能
  • 适合研究人体交互、智能视频分析的学者使用

人体-场景视觉语言任务在社会应用中日益普遍,但现有进展多依赖针对单一任务设计的模型。尽管大型视觉语言模型(VLMs)可提升多种下游任务表现,通用模型在专业领域仍表现不佳。本文提出领域专用大视觉语言模型HumanVLM,专为人体-场景任务设计。首先,构建大规模人体场景多模态图像-文本数据集HumanCaption-10M;其次,开发以人为中心的图像描述方法,构建高质量数据集HumanCaptionHQ(约31.1万对),全面捕捉人脸、身体及背景信息;最后,基于上述数据训练HumanVLM。实验表明,该模型在同类规模多模态模型中整体表现最优,尤其在人相关任务上显著优于Qwen2VL和ChatGPT-4o。HumanVLM及其数据集将推动人体周边领域的研究发展。

原文摘要 · Abstract (English)

Human-scene vision-language tasks are increasingly prevalent in diverse social applications, yet recent advancements predominantly rely on models specifically tailored to individual tasks. Emerging research indicates that large vision-language models (VLMs) can enhance performance across various downstream vision-language understanding tasks. However, general-domain models often underperform in specialized fields. This study introduces a domain-specific Large Vision-Language Model, Human-Scene Vision-Language Model (HumanVLM), designed to provide a foundation for human-scene Vision-Language tasks. Specifically, (1) we create a large-scale human-scene multimodal image-text dataset (HumanCaption-10M) sourced from the Internet to facilitate domain-specific alignment; (2) develop a captioning approach for human-centered images, capturing human faces, bodies, and backgrounds, and construct a high-quality Human-Scene image-text dataset (HumanCaptionHQ, about 311k pairs) that contain as much detailed information as possible about human; (3) Using HumanCaption-10M and HumanCaptionHQ, we train a HumanVLM. In the experiments, we then evaluate our HumanVLM across varous downstream tasks, where it demonstrates superior overall performance among multimodal models of comparable scale, particularly excelling in human-related tasks and significantly outperforming similar models, including Qwen2VL and ChatGPT-4o. HumanVLM, alongside the data introduced, will stimulate the research in human-around fields.

视觉语言模型人体识别多模态数据图像描述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。