无需标注即可恢复损坏视频,利用视觉大模型实现盲区修复。
Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
- 用视觉大模型引导,自动识别不同类型的码流损坏。
- 在无标注情况下,恢复质量显著优于现有方法。
- 适合需要高效修复的多媒体系统和实际应用。
视频信号在多媒体通信与存储中易受码流域损坏影响,轻微损坏即可能导致像素域严重退化。为从损坏输入中恢复真实时空内容,码流损坏视频恢复成为一项挑战性且研究不足的任务。现有方法需对每帧视频进行耗时的人工标注,工作量大;且因部分局部残差信息误导特征补全,高质量恢复困难。本文提出首个盲式码流损坏视频恢复框架,融合视觉基础模型与恢复模型,适配多种损坏类型及码流级提示。提出的检测任意损坏(DAC)模型结合视觉基础模型先验、码流与损坏知识,提升损坏定位与盲恢复能力。引入新型损伤感知特征补全(CFC)模块,基于高层损坏理解自适应处理残差贡献。通过视觉基础模型引导的分层特征增强与混合残差专家(MoRE)结构中的高层协调,有效抑制伪影并增强有用残差。全面评估表明,该方法在无需人工标注掩码序列下实现卓越性能,有助于提升用户体验、拓展应用场景,推动更可靠的多媒体通信与存储系统发展。
原文摘要 · Abstract (English)
Video signals are vulnerable in multimedia communication and storage systems, as even slight bitstream-domain corruption can lead to significant pixel-domain degradation. To recover faithful spatio-temporal content from corrupted inputs, bitstream-corrupted video recovery has recently emerged as a challenging and understudied task. However, existing methods require time-consuming and labor-intensive annotation of corrupted regions for each corrupted video frame, resulting in a large workload in practice. In addition, high-quality recovery remains difficult as part of the local residual information in corrupted frames may mislead feature completion and successive content recovery. In this paper, we propose the first blind bitstream-corrupted video recovery framework that integrates visual foundation models with a recovery model, which is adapted to different types of corruption and bitstream-level prompts. Within the framework, the proposed Detect Any Corruption (DAC) model leverages the rich priors of the visual foundation model while incorporating bitstream and corruption knowledge to enhance corruption localization and blind recovery. Additionally, we introduce a novel Corruption-aware Feature Completion (CFC) module, which adaptively processes residual contributions based on high-level corruption understanding. With VFM-guided hierarchical feature augmentation and high-level coordination in a mixture-of-residual-experts (MoRE) structure, our method suppresses artifacts and enhances informative residuals. Comprehensive evaluations show that the proposed method achieves outstanding performance in bitstream-corrupted video recovery without requiring a manually labeled mask sequence. The demonstrated effectiveness will help to realize improved user experience, wider application scenarios, and more reliable multimedia communication and storage systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。