无需重训练,加速交互式视频模型推理。
Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models

- 根据用户操作动态调整上下文与去噪计算。
- 在多个数据集上实现最高2.59倍加速。
- 适合需要实时交互的虚拟场景应用。
交互式视频世界模型可逐段生成视频以响应用户控制的摄像机移动,适用于实时游戏模拟、虚拟场景导航和具身智能训练。然而,由于上下文记忆增长、注意力复杂度呈二次增长以及重复去噪步骤,扩展到长交互轨迹成本极高。本文提出 Light Interaction,一种无需重训练的交互式视频世界模型推理加速框架。核心洞察是:交互自然支持轨迹依赖的自适应计算——在探索新区域时可丢弃空间记忆,根据局部潜在动态调整时间上下文,当摄像机返回熟悉区域时可复用早期模型输出。基于此,Light Interaction 结合自适应上下文管理、去噪缓存加速及软硬件协同设计的3D块稀疏注意力与融合Triton内核。在 HY-WorldPlay 与 Matrix-Game-3.0 上评估,Light Interaction 在不重新训练模型的情况下实现最高2.59倍速度提升,同时保持良好的视觉质量。
原文摘要 · Abstract (English)
Interactive video world models generate video chunk by chunk in response to user-controlled camera movements, enabling applications such as real-time game simulation, virtual scene navigation, and embodied AI training. However, scaling to long interactive trajectories is prohibitively expensive due to growing context memory, quadratic attention complexity, and repeated denoising steps. We present Light Interaction, a training-free inference acceleration framework for interactive video world models. Our key insight is that interaction naturally enables trajectory-dependent adaptive computation: retrieved spatial memory can be discarded during novel exploration, temporal context can be adjusted according to local latent dynamics, and early-step model outputs can be reused when the camera revisits familiar regions. Based on this insight, Light Interaction combines adaptive context management, denoising cache acceleration, and hardware-software co-designed 3D block sparse attention with fused Triton kernels. Evaluated on HY-WorldPlay and Matrix-Game-3.0, Light Interaction achieves up to 2.59x speedup without model retraining while maintaining competitive visual quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。