arXiv:2501.07713cs.CVcs.HC2025-01被引 4

对比工业场景下手部分割模型在分布内与分布外的性能差异

Testing Human-Hand Segmentation on In-Distribution and Out-of-Distribution Data in Human-Robot Interactions Using a Deep Ensemble Model

  • 采用深度集成模型融合UNet与RefineNet进行手部分割
  • 工业数据训练模型在分布外场景中表现更优,尤其面对罕见手势和模糊动作
  • 结合头戴与固定摄像头多视角数据,提升真实场景评估可靠性

可靠的人手检测与分割对提升人机协作安全性和促进高级交互至关重要。现有研究主要在分布内(ID)数据上评估手部分割性能,但无法覆盖真实场景中的分布外(OOD)情况。本文通过在工业场景下构建多样化数据集(含简单/复杂背景、工业工具、0至4只手、戴手套与否),模拟实际人机交互环境。针对分布外场景,引入手指交叉、快速运动导致的运动模糊等罕见条件,以应对认知不确定性和随机不确定性。采用头戴式相机与固定摄像头双视角采集RGB图像,评估基于现有头戴与固定摄像头数据集训练的模型表现。使用包含UNet与RefineNet的深度集成模型进行分割,并通过预测熵实现不确定性量化。结果表明:工业数据训练的模型显著优于非工业数据训练模型;尽管所有模型在分布外场景中均表现下降,但工业数据模型仍展现出更强泛化能力。

原文摘要 · Abstract (English)

Reliable detection and segmentation of human hands are critical for enhancing safety and facilitating advanced interactions in human-robot collaboration. Current research predominantly evaluates hand segmentation under in-distribution (ID) data, which reflects the training data of deep learning (DL) models. However, this approach fails to address out-of-distribution (OOD) scenarios that often arise in real-world human-robot interactions. In this study, we present a novel approach by evaluating the performance of pre-trained DL models under both ID data and more challenging OOD scenarios. To mimic realistic industrial scenarios, we designed a diverse dataset featuring simple and cluttered backgrounds with industrial tools, varying numbers of hands (0 to 4), and hands with and without gloves. For OOD scenarios, we incorporated unique and rare conditions such as finger-crossing gestures and motion blur from fast-moving hands, addressing both epistemic and aleatoric uncertainties. To ensure multiple point of views (PoVs), we utilized both egocentric cameras, mounted on the operator's head, and static cameras to capture RGB images of human-robot interactions. This approach allowed us to account for multiple camera perspectives while also evaluating the performance of models trained on existing egocentric datasets as well as static-camera datasets. For segmentation, we used a deep ensemble model composed of UNet and RefineNet as base learners. Performance evaluation was conducted using segmentation metrics and uncertainty quantification via predictive entropy. Results revealed that models trained on industrial datasets outperformed those trained on non-industrial datasets, highlighting the importance of context-specific training. Although all models struggled with OOD scenarios, those trained on industrial datasets demonstrated significantly better generalization.

手部分割工业机器人分布外深度集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。