arXiv:2511.22873cs.CVcs.IR2025-11中稿 · poster presentatio…

用远距离监控视频识别行人年龄性别,助力交通设计更安全

CNN-Based Framework for Pedestrian Age and Gender Classification Using Far-View Surveillance in Mixed-Traffic Intersections

  • 基于全身体态特征,用CNN实现六类行人分类
  • 最高准确率达86.19%,轻量模型训练更快参数更少
  • 无需人脸识别,适合现有摄像头实时部署

行人安全在交通混杂的都市路口仍是重大挑战,尤其在低收入和中等收入国家。年龄与性别显著影响行人脆弱性,但实时监测系统极少捕捉此类信息。本文提出一种基于卷积神经网络(CNN)的框架,从远距离路口视频中识别行人年龄组与性别,不依赖人脸识别或高分辨率图像。分类任务为统一的六类问题:成年、青少年、儿童男性与女性,依据全身视觉线索。数据采集自孟加拉国达卡三个高风险路口。采用两种CNN架构:预训练的ResNet50与定制轻量CNN。共测试八种模型变体,结合不同池化策略与优化器。结果显示,使用最大池化与SGD的ResNet50达到最高准确率86.19%;定制CNN表现接近(84.15%),参数更少,训练更快。模型设计高效,可在标准监控流上实现实时推理。对从业者而言,该系统可利用现有摄像头基础设施,低成本、规模化地监测路口行人构成,支持路口设计优化、信号配时调整及针对儿童、老人等弱势群体的精准干预。通过补充传统交通数据中缺失的人口统计信息,该框架助力混合交通环境下的包容性、数据驱动型规划。

原文摘要 · Abstract (English)

Pedestrian safety remains a pressing concern in congested urban intersections, particularly in low- and middle-income countries where traffic is multimodal, and infrastructure often lacks formal control. Demographic factors like age and gender significantly influence pedestrian vulnerability, yet real-time monitoring systems rarely capture this information. To address this gap, this study proposes a deep learning framework that classifies pedestrian age group and gender from far-view intersection footage using convolutional neural networks (CNNs), without relying on facial recognition or high-resolution imagery. The classification is structured as a unified six-class problem, distinguishing adult, teenager, and child pedestrians for both males and females, based on full-body visual cues. Video data was collected from three high-risk intersections in Dhaka, Bangladesh. Two CNN architectures were implemented: ResNet50, a deep convolutional neural network pretrained on ImageNet, and a custom lightweight CNN optimized for computational efficiency. Eight model variants explored combinations of pooling strategies and optimizers. ResNet50 with Max Pooling and SGD achieved the highest accuracy (86.19%), while the custom CNN performed comparably (84.15%) with fewer parameters and faster training. The model's efficient design enables real-time inference on standard surveillance feeds. For practitioners, this system provides a scalable, cost-effective tool to monitor pedestrian demographics at intersections using existing camera infrastructure. Its outputs can shape intersection design, optimize signal timing, and enable targeted safety interventions for vulnerable groups such as children or the elderly. By offering demographic insights often missing in conventional traffic data, the framework supports more inclusive, data-driven planning in mixed-traffic environments.

行人识别智能交通轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。