用扩散模型结合头骨结构引导,从X光片生成更匹配的面部图像。
Cranio-Diff: Diffusion-based Cross-domain Craniofacial Reconstruction with 2D X-ray Skull Guidance and Structural Identity Constraints

- 引入控制网和生物特征文本,实现头骨到面部的结构对齐生成
- 在4320组数据上验证,生成质量与检索性能优于现有方法
- 适用于法医重建,可为案件提供辅助视觉证据
当前先进的生成模型如CycleGAN、Pix2Pix和扩散模型在人脸生成任务中表现优异,但在从头骨(X光)到面部(光学)的跨模态颅面重建中,因模态间结构身份对齐不一致而难以有效捕捉语义信息。为此,我们提出Cranio-Diff,一种基于扩散模型的跨域颅面重建框架,通过控制网引入头骨条件结构引导,并结合生物特征文本条件,生成与给定头骨在语义和结构上更一致的面部图像。该方法在由120名受试者侧位与正位X光扫描构建的颅面数据集上评估,每张面部图像在三个年龄组(25、45、65岁)及三种体重指数变化(-10%、基准值、+10%)下合成,共生成4320对样本。据我们所知,这是唯一具备如此规模的X光-面部数据集。大量实验表明,所提方法在生成图像质量与检索任务中均优于现有方法。使用FID、IS、SSIM、LPIPS、PSNR和ArcFace分数评估生成质量,以recall@k、mAP@k和MRR@k评估检索性能。结果表明,该方法可作为法医调查中的替代工具。
原文摘要 · Abstract (English)
The state-of-the-art generative models, such as CycleGAN, Pix2Pix, and diffusion models have demonstrated remarkable performance in the face generation task. However, they fail to effectively capture cross-modality semantic information in craniofacial reconstruction when translating from the skull (x-ray) to the face (optical) domain, due to a mismatch in the alignment of structural identity across modalities. To address this issue, we propose Cranio-Diff, a diffusion-based framework for cross-domain cranio-facial reconstruction from 2D X-ray skull images. The proposed approach integrates skull-conditioned structural guidance through ControlNet with biometric text conditioning to generate a face which is more semantically and structurally aligned with the given skull. The proposed Cranio-diff method is evaluated on skull-face dataset obtained from X-ray scans of 120 subjects in lateral and frontal views. To enable controlled evaluation, each face image is synthesised across three age groups (25, 45, 65) and three BMI variations of -10%, baseline and +10%, yielding 4320 paired samples. To the best of our knowledge, this is the only X-ray-face dataset with this magnitude. Extensive experiments showed that the proposed method outperforms recent existing approaches in both generated image quality and retrieval task. Finally, to evaluate the performance of our proposed method, we have evaluated the quality of the generated image using FID, IS, SSIM, LPIPS, PSNR and ArcFace score. Additionally, retrieval performance is evaluated using recall@k, mAP@k and MRR@k. Obtained experimental results demonstrate that the proposed method can be used as an alternate tool in providing aid in forensic investigations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。