用改进GAN生成高质量印度手语图像,提升识别效果。
Generation of Indian Sign Language Letters, Numbers, and Words
- 融合ProGAN与SAGAN优点,设计带注意力机制的生成模型。
- 在印度手语字母、数字和129个词汇上实现高分辨率生成,IS提升3.2,FID降低30.12。
- 开源包含129个词的大规模高质量数据集,助力手语技术研究。
手语通过手势、面部表情和身体动作传达信息,是听障人士的重要沟通方式。尽管识别技术已取得进展,但手语生成仍需深入探索。本文提出一种结合渐进式生成对抗网络(ProGAN)与自注意力生成对抗网络(SAGAN)优势的改进GAN模型,用于生成高分辨率、特征丰富的类条件手语图像。该模型在生成印度手语字母、数字及129个词汇图像时表现优异,相较传统ProGAN,Inception Score(IS)提升3.2,Fréchet Inception Distance(FID)降低30.12。同时,我们发布了包含高质量印度手语字母、数字及129个词汇的大规模数据集,为相关研究提供支持。
原文摘要 · Abstract (English)
Sign language, which contains hand movements, facial expressions and bodily gestures, is a significant medium for communicating with hard-of-hearing people. A well-trained sign language community communicates easily, but those who don't know sign language face significant challenges. Recognition and generation are basic communication methods between hearing and hard-of-hearing individuals. Despite progress in recognition, sign language generation still needs to be explored. The Progressive Growing of Generative Adversarial Network (ProGAN) excels at producing high-quality images, while the Self-Attention Generative Adversarial Network (SAGAN) generates feature-rich images at medium resolutions. Balancing resolution and detail is crucial for sign language image generation. We are developing a Generative Adversarial Network (GAN) variant that combines both models to generate feature-rich, high-resolution, and class-conditional sign language images. Our modified Attention-based model generates high-quality images of Indian Sign Language letters, numbers, and words, outperforming the traditional ProGAN in Inception Score (IS) and Fréchet Inception Distance (FID), with improvements of 3.2 and 30.12, respectively. Additionally, we are publishing a large dataset incorporating high-quality images of Indian Sign Language alphabets, numbers, and 129 words.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。