arXiv:2604.02032cs.CVcs.LG2026-04中稿 · CVPR

构建首个多场景室内人群数据集,支持检测、分割与跟踪的自动标注。

IndoorCrowd: A Multi-Scene Dataset for Human Detection, Segmentation, and Tracking with an Automated Annotation Pipeline

  • 基于四所校园真实场景采集31段视频,含9913帧高精度人工标注掩码。
  • 自动标注工具在620帧测试集上表现接近人工,复杂场景中遮挡影响显著。
  • 提供可直接用于跟踪基准测试的2552帧连续身份轨迹,适合智能安防研究者。

理解拥挤室内环境中的人类行为对监控、智慧建筑和人机交互至关重要,但现有数据集很少以大规模捕捉真实室内复杂性。本文提出IndoorCrowd,一个涵盖四个校园地点(ACS-EC、ACS-EG、IE-Central、R-Central)的多场景数据集,用于室内人体检测、实例分割与多目标跟踪。数据集包含31段视频(共9,913帧,5帧/秒),每帧均有人工验证的逐实例分割掩码。其中620帧作为控制子集,用于评估SAM3、GroundingSAM、EfficientGroundingSAM三种基础模型自动标注器的表现,对比指标包括Cohen's κ、AP、精确率、召回率及掩码IoU。另有2,552帧子集以MOTChallenge格式支持多目标跟踪任务,提供连续身份追踪。采用YOLOv8n、YOLOv26n、RT-DETR-L搭配ByteTrack、BoT-SORT、OC-SORT建立检测、分割与跟踪基线。各场景分析显示难度差异显著,主要受人群密度、目标尺度与遮挡影响:其中ACS-EC场景有79.3%帧为密集人群,平均实例尺寸仅60.8像素,为最挑战场景。项目主页见https://sheepseb.github.io/IndoorCrowd/。

原文摘要 · Abstract (English)

Understanding human behaviour in crowded indoor environments is central to surveillance, smart buildings, and human-robot interaction, yet existing datasets rarely capture real-world indoor complexity at scale. We introduce IndoorCrowd, a multi-scene dataset for indoor human detection, instance segmentation, and multi-object tracking, collected across four campus locations (ACS-EC, ACS-EG, IE-Central, R-Central). It comprises $31$ videos ($9{,}913$ frames at $5$fps) with human-verified, per-instance segmentation masks. A $620$-frame control subset benchmarks three foundation-model auto-annotators: SAM3, GroundingSAM, and EfficientGroundingSAM, against human labels using Cohen's $κ$, AP, precision, recall, and mask IoU. A further $2{,}552$-frame subset supports multi-object tracking with continuous identity tracks in MOTChallenge format. We establish detection, segmentation, and tracking baselines using YOLOv8n, YOLOv26n, and RT-DETR-L paired with ByteTrack, BoT-SORT, and OC-SORT. Per-scene analysis reveals substantial difficulty variation driven by crowd density, scale, and occlusion: ACS-EC, with $79.3\%$ dense frames and a mean instance scale of $60.8$px, is the most challenging scene. The project page is available at https://sheepseb.github.io/IndoorCrowd/.

人体检测实例分割多目标跟踪数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。