通过让图像序列的神经轨迹变直,提升模型预测能力与鲁棒性。
Learning predictable and robust neural representations by straightening image sequences
- 设计自监督目标,显式鼓励神经表示在时间上走直线路径。
- 模型在合成视频上学习到可预测、分拆几何/光照/语义的嵌入表征。
- 对噪声和对抗攻击更鲁棒,可作为通用正则化项提升其他训练效果。
预测是所有生物体的基本能力,也被提出作为感知表征学习的目标。近期研究发现,在灵长类视觉系统中,预测能力得益于比初始光感受器编码更平直的时间轨迹,使线性外推成为可能。受此启发,我们提出一种自监督学习(SSL)目标,显式量化并促进轨迹直线化。我们在模拟自然视频常见特性的平滑合成图像序列上训练深度前馈神经网络,验证了该目标的有效性。所学模型生成的神经嵌入具有可预测性,且能分解对象的几何、光照和语义属性。相比以往优化随机增强不变性的SSL方法,这些表示对噪声和对抗攻击更具鲁棒性。此外,将该直线化目标作为正则化项,可迁移至其他训练流程,表明直线化原则在鲁棒无监督学习中具有广泛适用性。
原文摘要 · Abstract (English)
Prediction is a fundamental capability of all living organisms, and has been proposed as an objective for learning sensory representations. Recent work demonstrates that in primate visual systems, prediction is facilitated by neural representations that follow straighter temporal trajectories than their initial photoreceptor encoding, which allows for prediction by linear extrapolation. Inspired by these experimental findings, we develop a self-supervised learning (SSL) objective that explicitly quantifies and promotes straightening. We demonstrate the power of this objective in training deep feedforward neural networks on smoothly-rendered synthetic image sequences that mimic commonly-occurring properties of natural videos. The learned model contains neural embeddings that are predictive, but also factorize the geometric, photometric, and semantic attributes of objects. The representations also prove more robust to noise and adversarial attacks compared to previous SSL methods that optimize for invariance to random augmentations. Moreover, these beneficial properties can be transferred to other training procedures by using the straightening objective as a regularizer, suggesting a broader utility for straightening as a principle for robust unsupervised learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。