用实时视觉技术识别童工,提升监测准确率与效率。
Artificial Intelligence as a Tool for Combating Child Labour: A Real-Time Edge Vision Pipeline for Child Detection and Age Estimation
- 构建端侧实时多任务检测与追踪系统,融合人脸、年龄估计与身份重识别。
- 儿童年龄估计算法误差仅1.944年,较主流模型降低近20年。
- 在津巴布韦农场实测中检测率提升36倍,误报率降至1.8-3.9倍。
全球仍有约1.38亿儿童从事童工,现有监测依赖周期性入户调查,存在严重漏检。本文提出一种纯研究原型的实时边缘视觉管道,探索为童工监测与救济系统(CLMRS)提供持续存在的证据链。该管道集成多任务人体与人脸检测(基于CerberusDet框架的YOLO26x主干)、级联式年龄估计(MiVOLO v2结合0-12岁儿童专用模型)、ByteTrack追踪、ArcFace与DINOv2重识别,并通过轨迹级融合生成可审查的人体记录。检测器在人检测上将[email protected]从0.390提升至0.683;儿童专用模型在仅含儿童的验证集上达到1.944年平均绝对误差,远优于主流开源方案的18-23年误差。使用FP8 TensorRT编译实现1.77倍加速,仅增加0.002年误差,使系统在嵌入式设备上运行速度超2倍实时。在26.8小时模拟视频中,系统发现634名儿童候选者,较前代多出349人。进一步在津巴布韦农场开展17天无人值守实地测试(共3870万帧,6个摄像头),对比每日考勤表:软件调优使检测率提升36倍,通过同时性约束合并身份,将过报率从9.1倍降至1.8-3.9倍,且无真实误合并案例。论文还记录了训练与量化过程中的失败经验,强调数据保护与人工介入机制的必要性。
原文摘要 · Abstract (English)
An estimated 138 million children remain in child labour worldwide, and the monitoring systems used by affected sectors, built on periodic household visits and interviews, systematically under-detect them. We present a real-time computer-vision pipeline, built and operated solely as a research prototype, that studies the feasibility of giving Child Labour Monitoring and Remediation Systems (CLMRS) a continuous, presence-based evidence channel. The pipeline combines a multi-task person and face detector (YOLO26x backbone in the CerberusDet framework), cascaded age estimation pairing MiVOLO v2 with a child-specialist model for ages 0-12, ByteTrack tracking, ArcFace and DINOv2 re-identification, and track-level fusion producing reviewable per-person records. The detector raises person [email protected] from 0.390 to 0.683 over the previous-generation baseline; the child specialist reaches 1.944 years MAE on children-only validation, where widely used open-source stacks err by 18-23 years. FP8 TensorRT compilation yields a 1.77x speedup at +0.002 years MAE, bringing the pipeline above twice real-time on embedded hardware. On 26.8 hours of proxy video the system finds 634 unique child candidates versus 285 for its predecessor. We further report a seventeen-day unattended field pilot on a farm in Zimbabwe (38.7 million frames, six cameras) evaluated against a daily attendance register: software tuning improved detection yield 36-fold, and identity consolidation under a simultaneity veto cut over-reporting from 9.1x to 1.8-3.9x with zero proven-false merges. We document training and quantisation failures alongside successes, and the data-protection and human-in-the-loop safeguards such a system requires.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。