用扩散模型生成真实感防护装备,提升工业场景手物交互检测效果
GlovEgo-HOI: Bridging the Synthetic-to-Real Gap for Industrial Egocentric Human-Object Interaction Detection
- 结合合成数据与扩散模型,真实增强现实图像中的防护装备
- 在新数据集上实现91.3%的交互检测准确率,优于基线模型
- 适合工业安全、智能监控领域研究者参考
工业场景下的第一人称手物交互(EHOI)分析对安全生产至关重要,但受限于领域特定标注数据稀缺,模型泛化能力不足。本文提出一种数据生成框架,通过合成数据结合基于扩散模型的方法,将逼真的个人防护装备(PPE)注入真实图像中以增强数据多样性。构建了全新的工业级第一人称手物交互基准数据集GlovEgo-HOI,并设计GlovEgo-Net模型,集成手套头(Glove-Head)与关键点头(Keypoint-Head)模块,利用手部姿态信息提升交互检测性能。大量实验表明该方法有效,显著改善模型在真实场景下的表现。为推动后续研究,本文公开GlovEgo-HOI数据集、数据增强流程及预训练模型,项目地址见GitHub。
原文摘要 · Abstract (English)
Egocentric Human-Object Interaction (EHOI) analysis is crucial for industrial safety, yet the development of robust models is hindered by the scarcity of annotated domain-specific data. We address this challenge by introducing a data generation framework that combines synthetic data with a diffusion-based process to augment real-world images with realistic Personal Protective Equipment (PPE). We present GlovEgo-HOI, a new benchmark dataset for industrial EHOI, and GlovEgo-Net, a model integrating Glove-Head and Keypoint- Head modules to leverage hand pose information for enhanced interaction detection. Extensive experiments demonstrate the effectiveness of the proposed data generation framework and GlovEgo-Net. To foster further research, we release the GlovEgo-HOI dataset, augmentation pipeline, and pre-trained models at: GitHub project.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。