arXiv:2509.09111cs.CV2025-09

构建细粒度手机使用行为数据集,助力安全与注意力监控

FPI-Det: a face--phone Interaction Dataset for phone-use detection and understanding

  • 构建包含2.29万张图像的跨场景人脸-手机交互数据集
  • 在极端尺度变化和遮挡条件下,模型检测准确率低于60%
  • 适合研究移动设备行为理解、人机交互感知的学者使用

移动设备的广泛使用给视觉系统在安全监测、工作效能评估和注意力管理方面带来新挑战。判断用户是否在使用手机不仅需要物体识别,还需理解行为上下文,涉及在不同条件下对人脸、手部与设备间关系的推理。现有通用基准无法充分捕捉此类细粒度人机交互。为此,我们提出FPI-Det数据集,包含22,879张图像,涵盖工作、教育、交通及公共场景中人脸与手机的同步标注。数据集具有极端尺度变化、频繁遮挡和多样采集条件等特点。我们评估了代表性YOLO与DETR检测器,提供基线结果,并分析了在物体尺寸、遮挡程度和环境差异下的性能表现。源代码与数据集已公开于https://github.com/KvCgRv/FPI-Det。

原文摘要 · Abstract (English)

The widespread use of mobile devices has created new challenges for vision systems in safety monitoring, workplace productivity assessment, and attention management. Detecting whether a person is using a phone requires not only object recognition but also an understanding of behavioral context, which involves reasoning about the relationship between faces, hands, and devices under diverse conditions. Existing generic benchmarks do not fully capture such fine-grained human--device interactions. To address this gap, we introduce the FPI-Det, containing 22{,}879 images with synchronized annotations for faces and phones across workplace, education, transportation, and public scenarios. The dataset features extreme scale variation, frequent occlusions, and varied capture conditions. We evaluate representative YOLO and DETR detectors, providing baseline results and an analysis of performance across object sizes, occlusion levels, and environments. Source code and dataset is available at https://github.com/KvCgRv/FPI-Det.

行为识别数据集多模态感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。