提出可量化人体真实感的评估指标,解决生成图像中肢体异常等问题
BodyMetric: Evaluating the Realism of Human Bodies in Text-to-Image Generation
- 基于3D人体结构与文本描述,训练可学习的真实感评分模型
- 在新数据集BodyRealism上达到0.89相关性,显著优于现有通用评价指标
- 适用于大规模模型评测与图像排序,尤其适合关注人体生成质量的研究者
当前最先进的文生图模型在生成人体图像时仍存在肢体缺失、姿势失真、部位模糊等常见缺陷。现有评估主要依赖耗时的人工判断,难以规模化。本文提出BodyMetric,一种可学习的图像人体真实感评估指标,利用从图像推断出的3D人体表示和文本描述作为多模态信号,并在专家标注的体感真实度数据集BodyRealism上进行训练。消融实验验证了模型架构及引入3D人体先验的有效性。相比仅反映通用偏好性的现有指标,BodyMetric能更精准捕捉人体相关瑕疵。通过该指标,首次实现对多个文生图模型生成人体能力的大规模基准测试,并成功用于按真实感得分对生成图像进行排序。
原文摘要 · Abstract (English)
Accurately generating images of human bodies from text remains a challenging problem for state of the art text-to-image models. Commonly observed body-related artifacts include extra or missing limbs, unrealistic poses, blurred body parts, etc. Currently, evaluation of such artifacts relies heavily on time-consuming human judgments, limiting the ability to benchmark models at scale. We address this by proposing BodyMetric, a learnable metric that predicts body realism in images. BodyMetric is trained on realism labels and multi-modal signals including 3D body representations inferred from the input image, and textual descriptions. In order to facilitate this approach, we design an annotation pipeline to collect expert ratings on human body realism leading to a new dataset for this task, namely, BodyRealism. Ablation studies support our architectural choices for BodyMetric and the importance of leveraging a 3D human body prior in capturing body-related artifacts in 2D images. In comparison to concurrent metrics which evaluate general user preference in images, BodyMetric specifically reflects body-related artifacts. We demonstrate the utility of BodyMetric through applications that were previously infeasible at scale. In particular, we use BodyMetric to benchmark the generation ability of text-to-image models to produce realistic human bodies. We also demonstrate the effectiveness of BodyMetric in ranking generated images based on the predicted realism scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。