arXiv:2409.07271cs.CV2024-09被引 1

用扩散模型生成真实面部麻痹图像,填补数据稀缺空白

CFCPalsy: Facial Image Synthesis with Cross-Fusion Cycle Diffusion Model for Facial Paralysis Individuals

  • 提出跨融合循环扩散模型,融合多维度面部特征
  • 生成图像更逼真且保持身份一致性,优于现有方法
  • 适合医疗影像生成与面部康复算法训练者使用

当前面部麻痹诊断仍依赖临床医生的主观判断,存在评估不一致问题。自动化评估虽有潜力,但面部麻痹数据集稀缺限制了机器学习模型发展。为此,本文提出基于扩散模型的交叉融合循环麻痹表情生成模型(CFCPalsy),通过融合不同面部信息特征,增强面部区域的视觉细节与纹理表现,合成高保真度的面部麻痹图像,准确反映不同程度与类型的面瘫。在常用公开临床数据集上进行了定性与定量评估,结果表明该方法显著优于现有先进方法,生成图像更真实且身份一致性强。

原文摘要 · Abstract (English)

Currently, the diagnosis of facial paralysis remains a challenging task, often relying heavily on the subjective judgment and experience of clinicians, which can introduce variability and uncertainty in the assessment process. One promising application in real-life situations is the automatic estimation of facial paralysis. However, the scarcity of facial paralysis datasets limits the development of robust machine learning models for automated diagnosis and therapeutic interventions. To this end, this study aims to synthesize a high-quality facial paralysis dataset to address this gap, enabling more accurate and efficient algorithm training. Specifically, a novel Cross-Fusion Cycle Palsy Expression Generative Model (CFCPalsy) based on the diffusion model is proposed to combine different features of facial information and enhance the visual details of facial appearance and texture in facial regions, thus creating synthetic facial images that accurately represent various degrees and types of facial paralysis. We have qualitatively and quantitatively evaluated the proposed method on the commonly used public clinical datasets of facial paralysis to demonstrate its effectiveness. Experimental results indicate that the proposed method surpasses state-of-the-art methods, generating more realistic facial images and maintaining identity consistency.

面部生成扩散模型医疗影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。