通过语义与时空引导提升视频超分辨率的细节与连贯性
Semantic and Temporal Integration in Latent Diffusion Space for High-Fidelity Video Super-Resolution
- 在潜在扩散空间中融合语义与时空信息进行生成控制
- 显著提升细节恢复能力与帧间时序一致性
- 适合需要高保真视频重建的应用场景
视频超分辨率(VSR)模型近年来在增强低分辨率视频方面取得了显著进展。然而,由于生成过程控制不足,实现与低分辨率输入的高度保真对齐并维持帧间时序一致性仍是重大挑战。本文提出语义与时空引导视频超分辨率(SeTe-VSR),在潜在扩散空间中引入高层语义信息及时空联合建模,实现细节恢复与时序连贯性的均衡。该方法不仅保留高现实感视觉内容,还显著提升生成质量。大量实验表明,SeTe-VSR 在细节恢复和主观感知质量上均优于现有方法,展现出在复杂视频超分辨率任务中的有效性。
原文摘要 · Abstract (English)
Recent advancements in video super-resolution (VSR) models have demonstrated impressive results in enhancing low-resolution videos. However, due to limitations in adequately controlling the generation process, achieving high fidelity alignment with the low-resolution input while maintaining temporal consistency across frames remains a significant challenge. In this work, we propose Semantic and Temporal Guided Video Super-Resolution (SeTe-VSR), a novel approach that incorporates both semantic and temporal-spatio guidance in the latent diffusion space to address these challenges. By incorporating high-level semantic information and integrating spatial and temporal information, our approach achieves a seamless balance between recovering intricate details and ensuring temporal coherence. Our method not only preserves high-reality visual content but also significantly enhances fidelity. Extensive experiments demonstrate that SeTe-VSR outperforms existing methods in terms of detail recovery and perceptual quality, highlighting its effectiveness for complex video super-resolution tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。