用两步生成法让假口罩脸更真实,提升检测识别效果
Two-Step Data Augmentation for Masked Face Detection and Recognition: Turning Fake Masks to Real
- 先规则扭曲再GAN转换,生成更自然的戴口罩人脸
- 仅用1.9万张图训练,性能超越纯规则方法
- 适合资源有限却需快速落地的口罩识别研究
缺乏大规模带口罩人脸数据集制约了口罩人脸检测与识别的发展。本文提出一种两步生成式数据增强框架,结合规则掩码变形与无配对图像到图像的GAN翻译,生成超越规则叠加的逼真带口罩人脸样本。在约19,100张目标域图像(IAMGAN规模的3.8%)上训练,或加入跨领域预训练后达59,600张(11.8%),所提方法持续优于仅使用规则变形的方法,并取得与IAMGAN互补的结果,表明两步均具贡献。评估直接基于生成样本进行,为定性分析;未使用FID、KID等量化指标,因任何真实分布都可能不公平偏袒训练数据更接近的模型。引入非掩码保留损失以减少非掩码区域失真并稳定训练,通过随机噪声注入提升样本多样性。注:本文原为课程作业,在资源受限下完成;因奖学金中断转做兼职维持研究,受公司数据限制,学期中从医学影像转向口罩任务。在延迟算力与无AI辅助下,边上课边完成,期末提交至小型会议,获无修改接受。后续未向顶级会议投稿,因持续缺资金。识别与检测下游评估未在截止前完成。此说明回应后续比较与批评中未考虑实际条件的问题。
原文摘要 · Abstract (English)
The absence of large-scale masked face datasets challenges masked face detection and recognition. We propose a two-step generative data augmentation framework combining rule-based mask warping with unpaired image-to-image translation via GANs, producing masked face samples that go beyond rule-based overlays. Trained on about 19,100 images in the target domain (3.8% of IAMGAN's scale), or, including out-of-domain transfer pretraining, 59,600 and 11.8%, the proposed approach yields consistent improvements over rule-based warping alone and achieves results complementary to IAMGAN's, showing that both steps contribute. Evaluation is conducted directly on the generated samples and is qualitative; quantitative metrics like FID and KID were not applied as any real reference distribution would unfairly favor the model with closer training data. We introduce a non-mask preservation loss to reduce non-mask distortions and stabilize training, and stochastic noise injection to enhance sample diversity. Note: The paper originated as a coursework submission completed under resource constraints. Following an inexplicable scholarship termination, the author took on part-time employment to maintain research continuity, which led to a mid-semester domain pivot from medical imaging to masked face tasks due to company data restrictions. The work was completed alongside concurrent coursework with delayed compute access and without any AI assistance. It was submitted to a small venue at the semester end under an obligatory publication requirement and accepted without revision requests. Subsequent invitations to submit to first-tier venues were not pursued due to continued funding absence. Downstream evaluation on recognition or detection performance was not completed by the submission deadline. The note is added in response to subsequent comparisons and criticisms that did not account for these conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。