首个融合力觉与第一视角视频的大型数据集,助力物理动作理解
FEEL (Force-Enhanced Egocentric Learning): A Dataset for Physical Action Understanding
- 用自研压阻手套采集力觉信号,同步第一视角视频
- 300万帧自然操作数据,45%含手物接触,实现端到端接触分割
- 无需人工标注即可提升动作表征学习,适用于多场景迁移
我们提出FEEL(Force-Enhanced Egocentric Learning),首个将定制压阻手套采集的力觉信号与第一视角视频同步的大规模数据集。该手套支持可扩展的数据采集,FEEL包含约300万帧在厨房环境中自然、非脚本化的操作视频,其中45%的帧涉及手物接触。由于力是物理交互的根本驱动力,其在物理动作理解中具有关键作用。我们通过两个任务验证了力觉的有效性:(1) 接触理解,联合进行时间上接触分割与像素级被接触物体分割;(2) 动作表征学习,以力预测作为视频主干的自监督预训练目标。我们在时间接触分割上达到最先进性能,像素级分割表现良好,且无需人工标注接触物体分割。此外,基于FEEL的动作表征学习在无任何人工标签的情况下,显著提升了在EPIC-Kitchens、SomethingSomething-V2、EgoExo4D和Meccano上的迁移性能。
原文摘要 · Abstract (English)
We introduce FEEL (Force-Enhanced Egocentric Learning), the first large-scale dataset pairing force measurements gathered from custom piezoresistive gloves with egocentric video. Our gloves enable scalable data collection, and FEEL contains approximately 3 million force-synchronized frames of natural unscripted manipulation in kitchen environments, with 45% of frames involving hand-object contact. Because force is the underlying cause that drives physical interaction, it is a critical primitive for physical action understanding. We demonstrate the utility of force for physical action understanding through application of FEEL to two families of tasks: (1) contact understanding, where we jointly perform temporal contact segmentation and pixel-level contacted object segmentation; and, (2) action representation learning, where force prediction serves as a self-supervised pretraining objective for video backbones. We achieve state-of-the-art temporal contact segmentation results and competitive pixel-level segmentation results without any need for manual contacted object segmentation annotations. Furthermore we demonstrate that action representation learning with FEEL improves transfer performance on action understanding tasks without any manual labels over EPIC-Kitchens, SomethingSomething-V2, EgoExo4D and Meccano.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。