构建首个面向人类中心视觉理解与生成的统一基准,涵盖四类核心任务。
HUG-VIS: A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence

- 统一多模态数据集,含8400段同步视频、音频、文本与透明掩膜。
- 发现语言内容主导情绪识别,视觉情感识别最弱,边界保真度是关键瓶颈。
- 适合研究人机交互、跨模态生成及情感计算的学者与工程师使用。
视觉智能旨在感知、理解并合成视觉世界,是现代计算机视觉的核心。以人为中心的视觉智能尤为复杂,因其关注具表达力、社会情境化的人类主体,其意义常无法仅从外观推断。该领域融合视觉、语音与语言,覆盖四类代表性任务:人类情绪识别、人类视频生成、人类语音克隆和人类视频抠像。然而现有资源多为单任务设计,仅提供特定问题的模态与标注,缺乏协调理解与生成的共享基础,限制了多模态信号利用与研究拓展。为此,本文提出HUG-VIS——一个面向人类中心视觉理解与生成的统一基准。该数据集包含30位专业演员的8,400段坐姿半身视频,每人在受控普通话工作室协议下完成相同的280个情绪-动作-提示任务,配套同步的视频、音频、文本与透明掩膜(alpha mattes)。我们采用统一零样本协议,通过自动指标、任务特异性主观评分与多任务交叉分析,评估多种开源与闭源模型在四类任务上的表现。结果表明:(i) 当前情绪识别中语言内容起主导作用,纯视觉情感识别最弱;(ii) 视频生成与语音克隆中,自动指标与人工评价总体一致但排名差异显著,需联合报告;(iii) 运动中的边界保真度仍是视频抠像的主要挑战;(iv) 任务难度随情绪类型、模型与评估指标而异,且存在显著跨任务相关性。数据集与结果已公开于https://github.com/GML-MMGroup/HUG-VIS。
原文摘要 · Abstract (English)
Visual intelligence seeks to perceive, interpret, and synthesize the visual world and is central to modern computer vision. Human-centered visual intelligence is especially demanding because it studies people as expressive, socially situated subjects whose meaning is rarely conveyed by appearance alone. It couples vision with audio and language across four representative tasks: human emotion recognition, human video generation, human voice cloning, and human video matting. Yet existing resources remain task-specific, providing modalities and annotations for individual problems rather than a shared foundation coordinating understanding and generation. This limits multimodal signal use and broader research. We address this gap with HUG-VIS, a unified benchmark for Human-centered Understanding and Generation in Visual Intelligence. It contains 8,400 seated half-body videos of 30 professional actors, each performing the same 280 emotion-action-prompt assignments under a controlled Mandarin studio protocol, with synchronized video, audio, text, and alpha mattes. We evaluate diverse open- and closed-source models across the four tasks under a unified zero-shot protocol using automatic metrics, criterion-specific mean opinion scores, and multiple cross-task analyses. Results show that (i) linguistic content dominates current emotion recognition, while purely visual affect recognition is weakest; (ii) in video generation and voice cloning, automatic metrics and human judgment agree overall but differ in their top rankings, requiring joint reporting; (iii) boundary fidelity under motion is the main remaining obstacle for human matting; and (iv) task difficulty varies across emotions, models, and metrics, with notable cross-task correlations. The dataset and results are available at https://github.com/GML-MMGroup/HUG-VIS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。