arXiv:2411.14205cs.CVcs.AI2024-11CVPR被引 16

检测并修复生成人体图像中的结构异常,提升真实感。

Is this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-body

  • 提出细粒度人体异常检测框架,定位异常部位与类型。
  • 在两个高质量数据集上实现高精度检测,修复后视觉质量显著提升。
  • 适合需要真实人体图像的生成与质检场景,如影视、虚拟人。

近年来,视觉合成技术显著提升了生成人像的逼真度,但文本到图像或文本到视频模型常生成与现实人体结构差异较大的低质量人像,称为“异常人体”。这类异常难以检测与修复,需精准识别异常位置与类型。尽管视觉语言模型在多数视觉任务中表现优异,但在人体异常检测上效果不佳。本文首次提出细粒度人体异常检测任务(FHAD),构建两个高质量评估数据集,并提出名为HumanCalibrator的精细框架,可定位并修复人体结构异常,同时保留其他视觉内容。实验表明,该方法在异常检测上准确率高,修复后图像在视觉对比中表现更优,且有效保持原有内容。

原文摘要 · Abstract (English)

Recent improvements in visual synthesis have significantly enhanced the depiction of generated human photos, which are pivotal due to their wide applicability and demand. Nonetheless, the existing text-to-image or text-to-video models often generate low-quality human photos that might differ considerably from real-world body structures, referred to as "abnormal human bodies". Such abnormalities, typically deemed unacceptable, pose considerable challenges in the detection and repair of them within human photos. These challenges require precise abnormality recognition capabilities, which entail pinpointing both the location and the abnormality type. Intuitively, Visual Language Models (VLMs) that have obtained remarkable performance on various visual tasks are quite suitable for this task. However, their performance on abnormality detection in human photos is quite poor. Hence, it is quite important to highlight this task for the research community. In this paper, we first introduce a simple yet challenging task, i.e., \textbf{F}ine-grained \textbf{H}uman-body \textbf{A}bnormality \textbf{D}etection \textbf{(FHAD)}, and construct two high-quality datasets for evaluation. Then, we propose a meticulous framework, named HumanCalibrator, which identifies and repairs abnormalities in human body structures while preserving the other content. Experiments indicate that our HumanCalibrator achieves high accuracy in abnormality detection and accomplishes an increase in visual comparisons while preserving the other visual content.

人体生成异常检测图像修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。