arXiv:2606.07770cs.LG2026-06被引 1

对比学习易把慢变噪声当动态信号,新方法通过同轨迹采负样本解决此问题

Contrast encodes inductive bias: separating slow noise from dynamics in predictive representation learning

论文配图:Contrast encodes inductive bias: separating slow noise from dynamics in predictive representation learning
图 1 · 摘自论文原文
  • 同轨迹采负样本,阻断噪声伪装成动态的捷径
  • 长轨迹训练下即使强噪声也能学出更好表征
  • 适合物理系统建模、含噪声观测的动态学习场景

自监督方法如JEPA在潜在空间中学习表征并预测动态时,会将缓慢变化的噪声误认为真实动态信号。当噪声特征在每条轨迹内近似恒定,对比预测目标会优先编码这些特征而非系统的真实潜变量,导致表征被轨迹特异性噪声主导。下游性能随噪声强度下降,且增加训练轨迹数量和长度也无法改善。我们指出这一失败是对比目标固有缺陷,广泛存在于跨轨迹采负的对比预测方法中。在合成运动点数据集上的标准SimCLR式JEPA和刚体摆电影上的DySIB方法中,改用单轨迹内采负样本后,慢噪声无法区分帧内关系,消除了预测捷径。同时训练多个此类轨迹迫使编码器聚焦真正动态变量,更长轨迹仍能获得更优表征,即便在强慢噪声下亦然。结果揭示了动态表征学习中对比目标设计的关键原则,尤其适用于含噪声实验观测的物理系统。

原文摘要 · Abstract (English)

Self-supervised methods that learn representations and predict dynamics fully in the latent space, such as JEPA, have been shown to confuse slowly varying noise with the dynamical signals they aim to capture. Specifically, when noise features remain approximately constant within each trajectory, contrastive predictive objectives preferentially encode these features instead of the true latent variables governing the system. The learned representation then becomes dominated by trajectory-specific noise, so downstream performance degrades with noise strength and does not improve even as the number and duration of training trajectories increase. We argue that this failure is a property of the objective itself, shared by a long line of contrastive predictive objectives that sample negatives across trajectories. To illustrate this generality, we study the failure mode and its remedy in two settings: a standard SimCLR-style JEPA on a synthetic moving-dot dataset, and DySIB, a recently introduced method designed for extracting physically interpretable representations of dynamics, on movies of a rigid-body pendulum. When negatives are instead sampled within a single trajectory, the slow noise can no longer distinguish frames within that trajectory, removing the predictive shortcut. Training one encoder simultaneously on many such trajectories then forces it to encode the variables relevant for the dynamics, with longer trajectories yielding better representations even for strong slow noise. Our results point toward principles for designing contrastive predictive objectives in dynamical representation learning, especially for physical systems with noisy experimental observations.

表示学习对比学习动态建模噪声鲁棒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。