用变分生成模型随机编码,增强表情识别长尾数据。
Semantic Data Augmentation for Long-tailed Facial Expression Recognition
- 在VAE-GAN潜空间引入随机性生成新样本。
- 在RAF-DB上缓解长尾分布,提升小类别识别率。
- 适用于数据稀缺的视觉任务,如医疗与驾驶监控。
面部表情识别在社交机器人、健康护理、驾驶员疲劳监测等场景中具有广泛应用前景。尽管计算机视觉领域对此研究广泛,但真实场景中的表情识别仍具挑战性,部分原因在于数据集存在长尾分布。近年来,许多研究采用数据增强应对长尾识别问题。本文提出一种新的语义增强方法:通过在VAE-GAN的潜空间中对源数据编码引入随机性,生成新样本。针对RAF-DB数据集上的表情识别任务,该方法有效平衡了长尾分布。所提方法不仅适用于表情识别任务,还可推广至更多依赖大量数据的场景。
原文摘要 · Abstract (English)
Facial Expression Recognition has a wide application prospect in social robotics, health care, driver fatigue monitoring, and many other practical scenarios. Automatic recognition of facial expressions has been extensively studied by the Computer Vision research society. But Facial Expression Recognition in real-world is still a challenging task, partially due to the long-tailed distribution of the dataset. Many recent studies use data augmentation for Long-Tailed Recognition tasks. In this paper, we propose a novel semantic augmentation method. By introducing randomness into the encoding of the source data in the latent space of VAE-GAN, new samples are generated. Then, for facial expression recognition in RAF-DB dataset, we use our augmentation method to balance the long-tailed distribution. Our method can be used in not only FER tasks, but also more diverse data-hungry scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。