首个专用于评估生成图像中手部质量的评测体系,解决细节失真难题。
HandEval: Taking the First Step Towards Hand Quality Evaluation in Generated Images
- 基于成对图像构建无标注监督数据集,结合多模态大模型与手部关键点先验
- 在48k图像数据上训练,比现有方法更贴近人工判断
- 可提升生成手部真实感与AIGC检测精度,适合内容优化与安全检测场景
尽管近期文本到图像(T2I)模型显著提升了整体图像质量,但在复杂局部区域(如人体手部)的细节生成仍存在结构扭曲和纹理不真实等问题。当前手部质量评估几乎未受关注,制约了以人物为中心的生成质量优化与AIGC检测等下游任务。为此,本文提出首个面向生成手部区域的质量评估任务,并展示其丰富下游应用。我们首先构建了包含48,000张高质量与低质量手部成对图像的HandPair数据集,实现低成本、高效监督。基于此,设计了HandEval——一个专为手部质量评估优化的模型,融合多模态大语言模型(MLLM)的强大视觉理解能力与手部关键点先验知识,具备强感知能力。进一步构建了人工标注测试集,涵盖多种SOTA T2I模型生成的手部图像。实验表明,HandEval在评估一致性上优于现有方法。将其集成至图像生成与AIGC检测流程后,显著提升手部真实感与检测准确率,验证其通用有效性。代码与数据集将公开。
原文摘要 · Abstract (English)
Although recent text-to-image (T2I) models have significantly improved the overall visual quality of generated images, they still struggle in the generation of accurate details in complex local regions, especially human hands. Generated hands often exhibit structural distortions and unrealistic textures, which can be very noticeable even when the rest of the body is well-generated. However, the quality assessment of hand regions remains largely neglected, limiting downstream task performance like human-centric generation quality optimization and AIGC detection. To address this, we propose the first quality assessment task targeting generated hand regions and showcase its abundant downstream applications. We first introduce the HandPair dataset for training hand quality assessment models. It consists of 48k images formed by high- and low-quality hand pairs, enabling low-cost, efficient supervision without manual annotation. Based on it, we develop HandEval, a carefully designed hand-specific quality assessment model. It leverages the powerful visual understanding capability of Multimodal Large Language Model (MLLM) and incorporates prior knowledge of hand keypoints, gaining strong perception of hand quality. We further construct a human-annotated test set with hand images from various state-of-the-art (SOTA) T2I models to validate its quality evaluation capability. Results show that HandEval aligns better with human judgments than existing SOTA methods. Furthermore, we integrate HandEval into image generation and AIGC detection pipelines, prominently enhancing generated hand realism and detection accuracy, respectively, confirming its universal effectiveness in downstream applications. Code and dataset will be available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。