量化虚拟人生成中肤色还原偏差,发现深色皮肤误差更大
True to Tone? Quantifying Skin Tone Fidelity and Bias in Photographic-to-Virtual Human Pipelines
- 自动分析面部肤色与光照,统一评估虚拟人生成全流程
- 19,848次渲染显示深色皮肤肤色误差显著更高
- 无需训练、低算力,适合大规模公平性检测
真实还原面部肤色对虚拟人(VH)的现实感、身份一致性与公平性至关重要。然而多数可用的头像生成流程依赖未经色彩校准的图像输入,易引入不一致与偏见。本文提出一种全自动、可扩展的方法,系统评估VH生成流程中的肤色保真度。方法包含肤色与光照提取、纹理再着色、实时渲染及定量色彩分析。基于芝加哥人脸数据库(CFD)的图像,比较了基于面颊区域采样和全脸多维掩码两种肤色提取策略,并结合预训练的TRUST光照隔离框架(无需微调)。提取的肤色应用于MetaHuman纹理,在多种光照条件下渲染。通过CIELAB空间中的ΔE与个体表型角(ITA)客观评估肤色一致性。该方法无需人工干预,除预训练光照补偿模块外无训练环节,计算成本低,支持大规模评估。共生成并分析约19,848个渲染实例,结果表明不同肤色提取策略存在表型依赖性,且深色皮肤始终存在更高的色度误差。
原文摘要 · Abstract (English)
Accurate reproduction of facial skin tone is essential for realism, identity preservation, and fairness in Virtual Human (VH) rendering. However, most accessible avatar creation pipelines rely on photographic inputs that lack colorimetric calibration, which can introduce inconsistencies and bias. We propose a fully automatic and scalable methodology to systematically evaluate skin tone fidelity across the VH generation pipeline. Our approach defines a full workflow that integrates skin color and illumination extraction, texture recolorization, real-time rendering, and quantitative color analysis. Using facial images from the Chicago Face Database (CFD), we compare skin tone extraction strategies based on cheek-region sampling, following the literature, and multidimensional masking derived from full-face analysis. Additionally, we test both strategies with lighting isolation, using the pre-trained TRUST framework, employed without any training or optimization within our pipeline. Extracted skin tones are applied to MetaHuman textures and rendered under multiple lighting configurations. Skin tone consistency is evaluated objectively in the CIELAB color space using the $ΔE$ metric and the Individual Typology Angle (ITA). The proposed methodology operates without manual intervention and, with the exception of pre-trained illumination compensation modules, the pipeline does not include learning or training stages, enabling low computational cost and large-scale evaluation. Using this framework, we generate and analyze approximately 19,848 rendered instances. Our results show phenotype-dependent behavior of extraction strategies and consistently higher colorimetric errors for darker skin tones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。