arXiv:2501.15111cs.CV2025-01被引 63

首个面向人类中心场景的多模态大模型,能同时理解视觉与音频信息。

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

  • 构建240万条人类中心视频数据集,含1400万条指令和详细标注。
  • 三分支结构自适应融合特征,提升个体相关场景理解能力。
  • 支持情绪识别、表情描述等任务,适合人机交互与视频分析场景。

在以人为中心的场景中,同时理解视觉与听觉信息至关重要。尽管近期出现的多模态模型可处理多种模态,但普遍因缺乏大规模专用数据集和非针对性架构,在人类中心场景中表现不佳。本文提出HumanOmni,业界首个面向人类中心的多模态大语言模型。我们构建了一个包含超过240万条人类中心视频片段的数据集,配有详细描述和逾1400万条指令,有助于理解多样化的以人为中心场景。HumanOmni包含三个针对不同场景的专用分支,可根据用户指令自适应融合特征,显著提升个体相关场景的视觉理解能力。此外,模型整合音频特征,确保对环境及个体的全面理解。实验验证了HumanOmni在情感识别、面部表情描述、动作理解等多种任务中的先进性能。该模型将开源,以促进学术界与产业界的进一步发展与合作。

原文摘要 · Abstract (English)

In human-centric scenes, the ability to simultaneously understand visual and auditory information is crucial. While recent omni models can process multiple modalities, they generally lack effectiveness in human-centric scenes due to the absence of large-scale, specialized datasets and non-targeted architectures. In this work, we developed HumanOmni, the industry's first human-centric Omni-multimodal large language model. We constructed a dataset containing over 2.4 million human-centric video clips with detailed captions and more than 14 million instructions, facilitating the understanding of diverse human-centric scenes. HumanOmni includes three specialized branches for understanding different types of scenes. It adaptively fuses features from these branches based on user instructions, significantly enhancing visual understanding in scenes centered around individuals. Moreover, HumanOmni integrates audio features to ensure a comprehensive understanding of environments and individuals. Our experiments validate HumanOmni's advanced capabilities in handling human-centric scenes across a variety of tasks, including emotion recognition, facial expression description, and action understanding. Our model will be open-sourced to facilitate further development and collaboration within both academia and industry.

多模态视频理解人类中心大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。