arXiv:2609.06353cs.CV2026-09

构建儿童协作行为数据集,助力自然互动中社交参与分析。

ChildGaze: A Benchmark Dataset for Collaborative Behavior Understanding in Children

论文配图:ChildGaze: A Benchmark Dataset for Collaborative Behavior Understanding in Children
图 1 · 摘自论文原文
  • 基于儿童视频构建标注数据集,每帧独立标注儿童协作状态。
  • 标注一致性达93.16%(原始一致率),边界框平均重叠度0.808。
  • 支持教育研究与人机交互,适合儿童行为分析与视觉模型评估。

理解儿童在游戏和学习活动中的协作行为对分析社会参与、同伴互动、共享注意力和参与度至关重要。可靠识别这些线索可支持儿童发展、教育分析与以人为中心的计算机视觉研究。然而,仅估计儿童视线位置并不足以判断其是否积极参与共同活动。为此,我们提出ChildGaze,一个以儿童为中心的行为标注数据集,基于ChildPlay视频集合[1]构建。ChildGaze引入两种独立标注于每帧中每个儿童的协作与非协作标签。数据集提供儿童及成人面部、左手、右手的边界框,并在行、人物、帧三个层级组织标注。当前版本包含27个标注视频文件,共10,641帧,73,268条身体部位标注行。通过两名标注者对1,187帧进行独立标注,评估标注可靠性:协作标签达到93.16%原始一致率与Cohen's kappa 0.8631;边界框标注整体平均IoU为0.808。使用预训练ViT与Swin Transformer模型的基线实验显示,儿童级准确率最高达97.44%,帧级准确率最高达96.80%。结果表明,ChildGaze为自然情境下的儿童-成人及同伴互动协作行为研究提供了可靠基准。

原文摘要 · Abstract (English)

Understanding collaborative behavior in children is important for analyzing social participation, peer interaction, shared attention, and engagement during play and learning activities. Reliable recognition of these cues can support research in child development, educational analysis, and human-centered computer vision. However, estimating where a child is looking does not necessarily reveal whether the child is actively participating in a shared activity. To support this higher-level analysis, we introduce ChildGaze, a child-centered behavioral annotation dataset built on the ChildPlay video collection [1]. ChildGaze introduces two behavioral labels, collaborative and non-collaborative, assigned independently to each child within a frame. The dataset provides face, left-hand, and right-hand bounding boxes for children and adults and organizes the annotations at the row, person, and frame levels. The current release contains 27 annotated video files, 10,641 frames, and 73,268 body-part annotation rows. Annotation reliability was evaluated on 1,187 frames using independent annotations from two annotators. The collaboration labels achieved 93.16% raw agreement and a Cohen's kappa of 0.8631, while bounding-box annotations achieved an overall mean IoU of 0.808. Baseline experiments with pretrained ViT and Swin Transformer models achieved up to 97.44% child-person-level accuracy and 96.80% frame-level accuracy, respectively. These results show that ChildGaze provides a reliable benchmark for studying collaborative behavior in naturalistic child-adult and peer interactions.

儿童行为协作分析数据集视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。