arXiv:2604.06939cs.CV2026-04被引 5

解决自回归视频生成中语义遗忘、视觉漂移和可控性丢失问题

Grounded Forcing: Bridging Time-Independent Semantics and Proximal Dynamics in Autoregressive Video Synthesis

论文配图:Grounded Forcing: Bridging Time-Independent Semantics and Proximal Dynamics in Autoregressive Video Synthesis
图 1 · 摘自论文原文
  • 用双记忆键值缓存分离全局语义与局部时序动态
  • 通过双参考旋转位置编码抑制生成过程中的视觉漂移
  • 基于邻近权重更新缓存,实现提示切换时的平滑语义继承

自回归视频生成虽具备无限时长生成潜力,却受三大耦合挑战制约:上下文限制导致语义遗忘、位置外推引发视觉漂移、交互指令切换时可控性下降。现有方法多孤立处理,难以保障长期一致性。本文提出Grounded Forcing框架,通过三个互锁机制连接时间无关语义与邻近动态:首先,设计双记忆键值缓存(Dual Memory KV Cache),将局部时序动态与全局语义锚点解耦,确保长期语义连贯性与身份稳定性;其次,提出双参考旋转位置编码注入(Dual-Reference RoPE Injection),将位置嵌入限制在训练流形内,使全局语义对时间保持不变;第三,构建非对称邻近重缓存(Asymmetric Proximity Recache),通过邻近加权缓存更新,实现提示切换时的平滑语义传承。三者协同作用,在维持稳定语义核心的同时支持灵活局部动态。大量实验表明,该方法显著提升长程一致性和视觉稳定性,为交互式长视频生成奠定坚实基础。

原文摘要 · Abstract (English)

Autoregressive video synthesis offers a promising pathway for infinite-horizon generation but is fundamentally hindered by three intertwined challenges: semantic forgetting from context limitations, visual drift due to positional extrapolation, and controllability loss during interactive instruction switching. Current methods often tackle these issues in isolation, limiting long-term coherence. We introduce Grounded Forcing, a novel framework that bridges time-independent semantics and proximal dynamics through three interlocking mechanisms. First, to address semantic forgetting, we propose a Dual Memory KV Cache that decouples local temporal dynamics from global semantic anchors, ensuring long-term semantic coherence and identity stability. Second, to suppress visual drift, we design Dual-Reference RoPE Injection, which confines positional embeddings within the training manifold while rendering global semantics time-invariant. Third, to resolve controllability issues, we develop Asymmetric Proximity Recache, which facilitates smooth semantic inheritance during prompt transitions via proximity-weighted cache updates. These components operate synergistically to tether the generative process to stable semantic cores while accommodating flexible local dynamics. Extensive experiments demonstrate that Grounded Forcing significantly enhances long-range consistency and visual stability, establishing a robust foundation for interactive long-form video synthesis.

视频生成自回归模型语义稳定提示迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。