用物理力学原理解决视频语言模型的时序断裂问题
SLAP: The Semantic Least Action Principle for Variational Video-Language Modeling

- 将语义动态类比为经典力学,构建基于黎曼流形的视频轨迹模型
- 通过边界值问题求解实现长时序插值,避免物体消失和能量失稳
- 无需像素级生成即可保持对象持续性,适合长视频理解任务
在大型视频语言模型(LVLMs)时代,稀疏帧采样带来的计算必要性造成了根本性的“时间鸿沟”,使模型无法捕捉关键因果转变。现有依赖生成幻觉(如潜在扩散)或自回归外推的方法往往在长时程中丧失语义一致性,出现物体消失和能量不稳定性。我们提出从概率生成到变分力学的范式转变,引入语义最小作用量原理(SLAP)。通过建立经典力学与语义动态之间的严格同构关系,将潜空间视频轨迹建模为由语义拉格朗日函数决定的黎曼流形上的路径。将插值任务表述为边界值问题,通过离散欧拉-拉格朗日方程求解,自然实现对象持久性而无需像素级渲染。大量实验验证了所提方法的有效性。
原文摘要 · Abstract (English)
In the era of Large Video-Language Models (LVLMs), the computational necessity of sparse frame sampling creates a fundamental ``temporal gap'', rendering models blind to critical causal transitions. Existing solutions relying on generative hallucination (e.g., latent diffusion) or autoregressive extrapolation often fail to maintain semantic consistency over long horizons, suffering from object vanishing and energetic instability. We propose a paradigm shift from probabilistic generation to variational mechanics with the \textbf{Semantic Least Action Principle (SLAP)}. Drawing a rigorous isomorphism between classical mechanics and semantic dynamics, we model the latent video trajectory as a path on a Riemannian manifold governed by a Semantic Lagrangian. By formulating the interpolation task as a Boundary Value Problem (BVP) solved via the discrete Euler-Lagrange equations, SLAP naturally enforces object persistence without pixel-level rendering. Extensive experiments show the effectiveness of our proposed SLAP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。