arXiv:2503.00811cs.CV2025-03被引 3

首个系统评估生成图像中人体畸变的工具,精准识别肢体错位等缺陷。

Evaluating and Predicting Distorted Human Body Parts for Generated Images

  • 构建基于视觉变换器的ViT-HD模型,专精检测生成图像的人体畸变
  • 在Distortion-5K数据集上实现0.899的F1分数与0.831的交并比
  • 发现近半数生成图像存在人体畸变,助力提升文本生成图像质量

当前文本到图像(T2I)模型虽能生成高质量图像,但人体结构准确性仍存挑战。生成图像常出现肢体增多、手指缺失、肢体变形或融合等问题。现有评估指标如Inception Score(IS)和Fréchet Inception Distance(FID)缺乏检测此类畸变的细粒度能力,而人工偏好指标则侧重抽象质量而非解剖准确性。为此,我们建立了首个针对生成图像中人体畸变的标准,并推出包含4,700张标注图像的Distortion-5K数据集,涵盖多种风格与畸变类型。基于此数据集,提出面向畸变检测的ViT-HD模型,其在畸变定位任务上优于现有分割模型与视觉语言模型,达到0.899的F1分数和0.831的交并比。此外,构建包含500个以人物为中心提示的Human Distortion Benchmark,用于评估四种主流T2I模型,结果显示近50%的生成图像存在畸变。该工作首次系统化评估生成人体的解剖准确性,为提升T2I模型真实性和应用价值提供工具支持。相关数据集与训练好的ViT-HD将发布于GitHub:https://github.com/TheRoadQaQ/Predicting-Distortion。

原文摘要 · Abstract (English)

Recent advancements in text-to-image (T2I) models enable high-quality image synthesis, yet generating anatomically accurate human figures remains challenging. AI-generated images frequently exhibit distortions such as proliferated limbs, missing fingers, deformed extremities, or fused body parts. Existing evaluation metrics like Inception Score (IS) and Fréchet Inception Distance (FID) lack the granularity to detect these distortions, while human preference-based metrics focus on abstract quality assessments rather than anatomical fidelity. To address this gap, we establish the first standards for identifying human body distortions in AI-generated images and introduce Distortion-5K, a comprehensive dataset comprising 4,700 annotated images of normal and malformed human figures across diverse styles and distortion types. Based on this dataset, we propose ViT-HD, a Vision Transformer-based model tailored for detecting human body distortions in AI-generated images, which outperforms state-of-the-art segmentation models and visual language models, achieving an F1 score of 0.899 and IoU of 0.831 on distortion localization. Additionally, we construct the Human Distortion Benchmark with 500 human-centric prompts to evaluate four popular T2I models using trained ViT-HD, revealing that nearly 50\% of generated images contain distortions. This work pioneers a systematic approach to evaluating anatomical accuracy in AI-generated humans, offering tools to advance the fidelity of T2I models and their real-world applicability. The Distortion-5K dataset, trained ViT-HD will soon be released in our GitHub repository: \href{https://github.com/TheRoadQaQ/Predicting-Distortion}{https://github.com/TheRoadQaQ/Predicting-Distortion}.

人体生成畸变检测图像评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。