arXiv:2506.10689cs.CV2025-06被引 2

用多年龄任务检测未成年人,提升真实场景下的准确率与鲁棒性。

Underage Detection through a Multi-Task and MultiAge Approach for Screening Minors in Unconstrained Imagery

  • 设计多任务模型,聚焦12/15/18/21岁关键年龄点进行判别。
  • 在ASORES-39k上将误判率降至4.068年,1%假成年率下检测F2达0.857。
  • 针对低质量图像和极端姿态测试,仍保持近100%召回率,适合真实应用。

在非受限图像中精准自动筛查未成年人,需应对数据分布偏移与儿童样本不足问题。本文提出基于冻结FaRL视觉语言主干的多任务架构,结合轻量级MLP共享特征,包含一个年龄回归头和四个二分类未成年头(12、15、18、21岁)。通过α重加权焦点损失与年龄平衡采样缓解类别不平衡,剔除阈值附近模糊样本以消除年龄间隙。在新构建的总体未成年检测基准(303,000张训练图,110,000张测试图)上评估,定义了‘ASORES-39k’(去噪域)与‘ASWIFT-20k’(极端姿态、表情、低画质)测试集。在清洗后的整体数据集上训练并采用重采样与年龄间隙策略,模型‘F’将ASORES-39k上的平均绝对误差从4.175年降至4.068年,1%假成年率下未成人检测F2得分由0.801提升至0.857;在ASWIFT-20k上仍维持接近0.99的召回率,F2从0.742升至0.833,展现强域偏移鲁棒性。

原文摘要 · Abstract (English)

Accurate automatic screening of minors in unconstrained images requires models robust to distribution shift and resilient to the under-representation of children in public datasets. To address these issues, we propose a multi-task architecture with dedicated under/over-age discrimination tasks based on a frozen FaRL vision-language backbone joined with a compact two-layer MLP that shares features across one age-regression head and four binary underage heads (12, 15, 18, and 21 years). This design focuses on the legally critical age range while keeping the backbone frozen. Class imbalance is mitigated through an $α$-reweighted focal loss and age-balanced mini-batch sampling, while an age gap removes ambiguous samples near thresholds. Evaluation is conducted on our new Overall Underage Benchmark (303k cleaned training images, 110k test images), defining both the "ASORES-39k" restricted overall test, which removes the noisiest domains, and the age estimation wild-shifts test "ASWIFT-20k" of 20k-images, stressing extreme poses ($>$45°), expressions, and low image quality to emulate real-world shifts. Trained on the cleaned overall set with resampling and age gap, our multiage model "F" reduces the mean absolute error on ASORES-39k from 4.175 y (age-only baseline) to 4.068 y and improves under-18 detection from F2 score of 0.801 to 0.857 at 1% false-adult rate. Under the ASWIFT-20k, the same configuration nearly sustains 0.99 recall while F2 rises from 0.742 to 0.833, demonstrating robustness to domain shift.

未成年人检测多任务学习域适应视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。