arXiv:2505.05531cs.CVeess.IV2025-05被引 1

用多维注意力UNet提升唇部分割精度,尤其改善了杯状弓区域的边界识别。

OXSeg: Multidimensional attention UNet-based lip segmentation using semi-supervised lip contours

  • 融合局部二值模式构建多维输入,通过序列注意力UNet重建唇轮廓。
  • 上唇分割平均Dice系数达84.75%,像素准确率99.77%,显著提升边界精度。
  • 适用于高精度唇形分析,特别适合胎儿酒精综合征等医学表型研究。

唇部分割在唇语识别、唇音同步和诊断等领域具有重要意义。然而,监督式分割受限于训练阶段唇轮廓标注的可用性,且对图像质量、光照和肤色敏感,导致边界检测不准确。为此,我们提出一种基于注意力UNet与多维输入的分步唇部分割方法。通过局部二值模式(LBP)挖掘面部图像的微结构特征,构建多维输入,并输入序列注意力UNet进行唇轮廓重建。设计了一种基于少数解剖标志点生成完整唇轮廓掩码的方法,在训练中提升分割精度。实验使用面部图像分割上唇,并评估胎儿酒精综合征(FAS)患者的唇部异常特征。所提方法在上唇分割中达到84.75%的平均Dice分数和99.77%的像素准确率。进一步结合生成对抗网络(GAN)分类器,对特定人群的FAS识别准确率达98.55%。该方法显著提升了杯状弓区域的分割精度,有助于揭示FAS特有的唇部形态特征。

原文摘要 · Abstract (English)

Lip segmentation plays a crucial role in various domains, such as lip synchronization, lipreading, and diagnostics. However, the effectiveness of supervised lip segmentation is constrained by the availability of lip contour in the training phase. A further challenge with lip segmentation is its reliance on image quality , lighting, and skin tone, leading to inaccuracies in the detected boundaries. To address these challenges, we propose a sequential lip segmentation method that integrates attention UNet and multidimensional input. We unravel the micro-patterns in facial images using local binary patterns to build multidimensional inputs. Subsequently, the multidimensional inputs are fed into sequential attention UNets, where the lip contour is reconstructed. We introduce a mask generation method that uses a few anatomical landmarks and estimates the complete lip contour to improve segmentation accuracy. This mask has been utilized in the training phase for lip segmentation. To evaluate the proposed method, we use facial images to segment the upper lips and subsequently assess lip-related facial anomalies in subjects with fetal alcohol syndrome (FAS). Using the proposed lip segmentation method, we achieved a mean dice score of 84.75%, and a mean pixel accuracy of 99.77% in upper lip segmentation. To further evaluate the method, we implemented classifiers to identify those with FAS. Using a generative adversarial network (GAN), we reached an accuracy of 98.55% in identifying FAS in one of the study populations. This method could be used to improve lip segmentation accuracy, especially around Cupid's bow, and shed light on distinct lip-related characteristics of FAS.

唇部分割注意力机制医学影像多模态输入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。