arXiv:2601.06212cs.CVcs.AI2026-01

用物理守恒定律提升多模态模型的视频生成速度与一致性

Akasha 2: Hamiltonian State Space Duality and Visual-Language Joint Embedding Predictive Architectur

  • 引入哈密顿状态空间对偶与稀疏哈密顿专家,让模型学习物理规律
  • 生成速度比扩散模型快4倍,延迟低于50毫秒,保持能量守恒
  • 适合需要实时视频生成和物理准确性的应用,如机器人控制

我们提出 Akasha 2,一种融合哈密顿状态空间对偶(H-SSD)与视觉语言联合嵌入预测架构(VL-JEPA)的前沿多模态模型。该系统采用 Mamba-3 选择性状态空间模型,结合稀疏哈密顿专家混合(SMoE-HE),通过辛积分强制潜在空间的物理守恒律。在视觉合成方面,引入哈密顿流匹配(HFM)与持续3D高斯溅射(3DGS),实现在移动设备上超低延迟(<50ms)生成。本工作建立了一种新的潜在世界模型范式,通过全息记忆架构实现前所未有的时空一致性。实验表明,将物理启发的归纳偏置融入神经架构可显著提升性能:视频预测达到当前最优水平(FVD: 287),视觉合成速度为扩散模型的4倍,推理速度比变压器基线快3-18倍,且在长时序下仍保持能量守恒。

原文摘要 · Abstract (English)

We present Akasha 2, a state-of-the-art multimodal architecture that integrates Hamiltonian State Space Duality (H-SSD) with Visual-Language Joint Embedding Predictive Architecture (VL-JEPA). The system leverages the Mamba-3 Selective State Space Model (SSM) augmented by a Sparse Mixture of Hamiltonian Experts (SMoE-HE) that enforces latent physical conservation laws through symplectic integration. For visual synthesis, we introduce Hamiltonian Flow Matching (HFM) and persistent 3D Gaussian Splatting (3DGS), enabling ultra-low latency (<50ms) on mobile hardware. This work establishes a new paradigm in latent world models, achieving unprecedented spatiotemporal coherence through a holographic memory architecture. Our approach demonstrates that incorporating physics-inspired inductive biases into neural architectures yields significant improvements: state-of-the-art video prediction (FVD: 287), 4x faster visual synthesis than diffusion models, and 3-18x inference speedup over transformer baselines while maintaining energy conservation over extended horizons.

多模态视频生成物理模型高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。