构建首个面向人脸生成的偏好评估基准,解决真实感与身份一致性的评测难题。
F-Bench: Rethinking Human Preference Evaluation Metrics for Benchmarking Face Generation, Customization, and Restoration
- 构建包含32742条评分的FaceQ数据库,覆盖29个模型在三类任务中的表现。
- 发现现有质量评估指标在真实性、身份保真度上效果不佳。
- 适合关注生成人脸可信度与可控性的研究者和开发者使用。
人工智能生成模型在人脸图像生成、定制和修复方面表现出强大能力,但常因独特失真、不自然细节和意外身份变化而偏离人类偏好,亟需全面的质量评估框架。为此,我们提出FaceQ,一个大规模、细粒度标注的AI生成人脸数据库,涵盖12,255张由29个模型生成的图像,涉及人脸生成、定制与修复三类任务。数据集包含180名标注者提供的32,742条平均意见分数(MOS),从质量、真实性、身份(ID)保真度及文本-图像一致性等多维度进行评估。基于FaceQ,我们建立F-Bench基准,用于系统比较和评估各类模型在不同提示下的表现。此外,我们评估了现有图像质量评估(IQA)、人脸质量评估(FQA)、AI生成内容图像质量评估(AIGCIQA)及偏好评估指标,发现它们在真实性、身份保真度和文本-图像一致性方面表现有限。FaceQ数据库将在论文发表后公开。
原文摘要 · Abstract (English)
Artificial intelligence generative models exhibit remarkable capabilities in content creation, particularly in face image generation, customization, and restoration. However, current AI-generated faces (AIGFs) often fall short of human preferences due to unique distortions, unrealistic details, and unexpected identity shifts, underscoring the need for a comprehensive quality evaluation framework for AIGFs. To address this need, we introduce FaceQ, a large-scale, comprehensive database of AI-generated Face images with fine-grained Quality annotations reflecting human preferences. The FaceQ database comprises 12,255 images generated by 29 models across three tasks: (1) face generation, (2) face customization, and (3) face restoration. It includes 32,742 mean opinion scores (MOSs) from 180 annotators, assessed across multiple dimensions: quality, authenticity, identity (ID) fidelity, and text-image correspondence. Using the FaceQ database, we establish F-Bench, a benchmark for comparing and evaluating face generation, customization, and restoration models, highlighting strengths and weaknesses across various prompts and evaluation dimensions. Additionally, we assess the performance of existing image quality assessment (IQA), face quality assessment (FQA), AI-generated content image quality assessment (AIGCIQA), and preference evaluation metrics, manifesting that these standard metrics are relatively ineffective in evaluating authenticity, ID fidelity, and text-image correspondence. The FaceQ database will be publicly available upon publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。