用合成数据提升第一人称视角下手物交互检测,尤其在真实标注数据少时效果显著。
Leveraging Synthetic Data for Enhancing Egocentric Hand-Object Interaction Detection
- 构建合成数据生成流水线,自动标注接触状态、边界框和像素级掩码。
- 仅用10%真实数据+合成数据,在三个数据集上提升5.67%~11.69%的检测准确率。
- 合成数据与真实场景越匹配,提升效果越明显,适合数据稀缺场景研究者。
本文探讨了合成数据在提升第一人称视角手物交互(HOI)检测中的作用。在VISOR、EgoHOS和ENIGMA-51数据集上通过大量实验与对比分析发现,当真实标注数据稀缺或缺失时,合成数据能显著提升检测性能。仅使用10%真实标注数据配合合成数据,模型在总体平均精度(Overall AP)上分别在VISOR、EgoHOS和ENIGMA-51上提升+5.67%、+8.24%和+11.69%。此外,系统研究了合成数据在物体、抓握方式和环境方面与真实数据对齐的影响,结果表明合成数据与真实数据的对齐程度越高,性能提升越显著。本工作发布了新的数据生成流程及HOI-Synth基准数据集,为现有数据集补充了带自动标注的合成手物交互图像。所有数据、代码与工具已公开:https://fpv-iplab.github.io/HOI-Synth/
原文摘要 · Abstract (English)
In this work, we explore the role of synthetic data in improving the detection of Hand-Object Interactions from egocentric images. Through extensive experimentation and comparative analysis on VISOR, EgoHOS, and ENIGMA-51 datasets, our findings demonstrate the potential of synthetic data to significantly improve HOI detection, particularly when real labeled data are scarce or unavailable. By using synthetic data and only 10% of the real labeled data, we achieve improvements in Overall AP over models trained exclusively on real data, with gains of +5.67% on VISOR, +8.24% on EgoHOS, and +11.69% on ENIGMA-51. Furthermore, we systematically study how aligning synthetic data to specific real-world benchmarks with respect to objects, grasps, and environments, showing that the effectiveness of synthetic data consistently improves with better synthetic-real alignment. As a result of this work, we release a new data generation pipeline and the new HOI-Synth benchmark, which augments existing datasets with synthetic images of hand-object interaction. These data are automatically annotated with hand-object contact states, bounding boxes, and pixel-wise segmentation masks. All data, code, and tools for synthetic data generation are available at: https://fpv-iplab.github.io/HOI-Synth/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。