用时间分段模拟记忆,提升第一视角视频的自监督学习效果。
Memory Storyboard: Leveraging Temporal Segmentation for Streaming Self-Supervised Learning from Egocentric Videos
- 将近期帧按时间分段构建记忆故事板,增强视觉流摘要能力。
- 在SAYCam和KrishnaCam数据集上优于当前最优无监督持续学习方法。
- 适合研究连续学习与第一视角视频表征的学者使用。
自监督学习有望从真实世界的连续非标注数据流中学习有效表征。然而,现有视觉自监督学习大多聚焦静态图像或人工数据流。为探索更真实的训练场景,我们研究了从长时真实第一视角视频流中进行流式自监督学习。受人类感知与记忆中事件分割机制启发,提出「记忆故事板」(Memory Storyboard)方法,将近期帧分组为时间片段,以更有效地总结历史视觉流用于记忆回放。为支持高效的时间分段,设计两级记忆架构:近期内容存于短期记忆,故事板时间片段则转移至长期记忆。在SAYCam和KrishnaCam等真实第一视角视频数据集上的实验表明,基于故事板帧的对比学习可生成语义有意义的表征,性能优于当前最先进的无监督持续学习方法。
原文摘要 · Abstract (English)
Self-supervised learning holds the promise of learning good representations from real-world continuous uncurated data streams. However, most existing works in visual self-supervised learning focus on static images or artificial data streams. Towards exploring a more realistic learning substrate, we investigate streaming self-supervised learning from long-form real-world egocentric video streams. Inspired by the event segmentation mechanism in human perception and memory, we propose "Memory Storyboard" that groups recent past frames into temporal segments for more effective summarization of the past visual streams for memory replay. To accommodate efficient temporal segmentation, we propose a two-tier memory hierarchy: the recent past is stored in a short-term memory, and the storyboard temporal segments are then transferred to a long-term memory. Experiments on real-world egocentric video datasets including SAYCam and KrishnaCam show that contrastive learning objectives on top of storyboard frames result in semantically meaningful representations that outperform those produced by state-of-the-art unsupervised continual learning methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。