arXiv:2506.20879cs.CV2025-06NeurIPS被引 6

构建首个多人体图像生成基准,评估身份一致性与动作准确性。

MultiHuman-Testbench: Benchmarking Image Generation for Multiple Humans

  • 设计包含1800条提示的多人体生成评测集,匹配5550张人脸图像。
  • 提出四项指标评估人脸数量、身份相似度、提示对齐和动作识别准确率。
  • 引入人体分割与匈牙利匹配技术,显著提升多人身份保持能力。

生成包含多个执行复杂动作的人体图像并保持其面部身份是重大挑战,主要源于缺乏专用评测基准。为此,我们提出 MultiHuman-Testbench,一个用于严格评估多人体生成模型的新基准。该基准包含1,800个精心设计的文本提示,涵盖从简单到复杂的多种人体动作;每个提示匹配5,550张独特的人脸图像,均匀采样以确保年龄、种族和性别多样性。同时提供由人工选择的姿势条件图像,精确对应提示内容。我们设计了多维度评估体系,采用四项关键指标量化人脸数量、身份相似度、提示对齐程度及动作检测效果。对多种模型(包括零样本方法与基于训练的方法,有无区域先验)进行了全面评估,并提出利用人体分割与匈牙利匹配实现图像与区域隔离的新技术,显著提升身份保留效果。本研究为多人体图像生成领域提供了重要工具与洞察。数据集与评估代码将开源于 https://github.com/Qualcomm-AI-research/MultiHuman-Testbench。

原文摘要 · Abstract (English)

Generation of images containing multiple humans, performing complex actions, while preserving their facial identities, is a significant challenge. A major factor contributing to this is the lack of a dedicated benchmark. To address this, we introduce MultiHuman-Testbench, a novel benchmark for rigorously evaluating generative models for multi-human generation. The benchmark comprises 1,800 samples, including carefully curated text prompts, describing a range of simple to complex human actions. These prompts are matched with a total of 5,550 unique human face images, sampled uniformly to ensure diversity across age, ethnic background, and gender. Alongside captions, we provide human-selected pose conditioning images which accurately match the prompt. We propose a multi-faceted evaluation suite employing four key metrics to quantify face count, ID similarity, prompt alignment, and action detection. We conduct a thorough evaluation of a diverse set of models, including zero-shot approaches and training-based methods, with and without regional priors. We also propose novel techniques to incorporate image and region isolation using human segmentation and Hungarian matching, significantly improving ID similarity. Our proposed benchmark and key findings provide valuable insights and a standardized tool for advancing research in multi-human image generation. The dataset and evaluation codes will be available at https://github.com/Qualcomm-AI-research/MultiHuman-Testbench.

图像生成多人体身份保持评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。