用稀疏注意力建模运动,实现流畅高清视频重建
MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling

- 基于世界模型思想,通过运动感知预测潜在状态
- 在保持时序一致性的同时,实现高质量视频重建
- 支持用户控制,平衡清晰度与平滑度,适合实时应用
视频超分辨率(VSR)旨在从低分辨率输入恢复高保真、高分辨率视频,广泛应用于移动拍摄、流媒体和档案修复。现有方法在局部细节保真度、长程时空建模、感知真实性和效率之间存在权衡:卷积对齐技术虽保留局部结构,但在大运动或复杂退化下表现不佳;基于Transformer的方法能捕捉长程依赖,但需架构或算法调整以保证计算可行性;近期的潜在空间或扩散生成器可合成丰富纹理,但需特殊时间约束以维持连贯性。我们提出MotionCraft,一种可控的VSR框架,将修复任务建模为受运动引导的潜在状态预测,融合自适应稀疏注意力与用户可访问的控制接口。该框架结合稳健的运动融合、兼顾局部与定向非局部交互的潜在世界变压器,以及紧凑的条件解码器,在流式传输约束下实现时序一致的高质量重建。实证评估表明,MotionCraft在重建和感知性能上表现优异,同时可预测地调节时序平滑度与重建保真度之间的权衡。
原文摘要 · Abstract (English)
Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from mobile capture to streaming and archival restoration. Existing approaches trade off among local-detail fidelity, long-range spatio-temporal modeling, perceptual realism, and efficiency: convolutional alignment techniques preserve local structure but suffer when motion is large or degradations are complex; transformer-based methods capture long-range dependencies yet require architectural or algorithmic adaptations to remain computationally feasible; and recent latent or diffusion-based generators synthesize rich texture but require specialized temporal constraints to maintain coherence. We present MotionCraft, a controllable VSR framework that formulates restoration as motion-aware latent state prediction inspired by world models and integrates adaptive sparse attention with an explicit user-accessible control interface. MotionCraft combines robust motion fusion, a Latent World Transformer that balances locality and targeted non-local interactions, and a compact conditional decoder to deliver temporally consistent, high-quality reconstructions under streaming constraints. Empirical evaluations show that MotionCraft achieves strong reconstruction and perceptual performance while enabling predictable trade-offs between temporal smoothness and reconstruction fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。