构建儿童危险行为识别数据集,提升安全监控能力
KidRisk: Benchmark Dataset for Children Dangerous Action Recognition

- 构建包含2500段视频、10000张图像的儿童危险行为数据集
- 视觉语言模型在危险动作识别上准确率达96.14%,显著优于传统方法
- 适合儿童安全监护、智能视频分析领域研究者参考
儿童天性活泼,在自主活动中常面临潜在危险,尤其在缺乏家长监督时。准确识别高风险动作对保障其安全至关重要。本文构建了一个新挑战性数据集KidRisk,包含2500段儿童行为短视频和10000张儿童危险动作图像。我们在该数据集上建立了基准测试,发现传统深度学习模型在此任务上表现有限。为此,我们提出了基于视觉-语言的基线模型,具备更强的上下文理解能力。所提方法在儿童行为分类上达到83.53%准确率,在危险动作识别上达96.14%,显著优于传统方法。结果表明,视觉-语言模型不仅可行,且在检测危险行为方面极具有效性,有助于提升儿童安全保障水平。
原文摘要 · Abstract (English)
Children are naturally energetic, and during their spontaneous activities, they often encounter potentially dangerous situations, especially when lacking parental supervision. Identifying actions that pose risks plays a crucial role in ensuring their safety. This paper build a novel challenging dataset, namely KidRisk, including 2,500 short videos of children's actions and 10,000 images for dangerous action of children. We also introduce a benchmark on our newly constructs dataset and find that traditional deep learning models demonstrated limited effectiveness on these tasks. Therefore, we develop vision-language based baselines with exceptional context understanding of visual information. Our proposed methods achieved an accuracy of 83.53% in classifying children's actions and 96.14% in recognizing children's dangerous actions, significantly outperforming traditional approaches. These results confirm that vision-language models are not only feasible but also highly effective in detecting hazardous actions, contributing positively to safeguarding children's safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。