用文字提示提升戏曲视频超分辨率,修复画质同时保留细节。
TextOVSR: Text-Guided Real-World Opera Video Super-Resolution
- 引入描述退化的文本和内容的文本双分支引导重建。
- 在自建戏曲低质数据集上,性能超越现有方法。
- 适合需要高质量戏曲视频修复的研究者与影视从业者。
许多经典戏曲视频因早期拍摄设备限制及长期存储导致画质下降。尽管真实世界视频超分辨率(RWVSR)近年取得进展,但直接应用于退化戏曲视频仍面临挑战。一是难以准确建模真实退化:简单组合传统退化核无法捕捉真实噪声分布,而从外部数据集提取噪声样本易引发风格不匹配并引入视觉伪影;二是现有方法仅依赖低质量图像特征,缺乏高层语义引导,难以重建真实细腻纹理。为此,本文提出文本引导的双分支戏曲视频超分辨率(TextOVSR)网络,引入两类文本提示:基于退化过程的描述性文本用于负分支以约束解空间,内容描述性文本结合提出的文本增强判别器(TED)提供语义引导以增强纹理重建。此外,设计退化鲁棒特征融合(DRF)模块,在抑制退化干扰的同时实现跨模态特征融合。在自建的OperaLQ基准测试中,TextOVSR在定性和定量上均优于当前最佳方法。代码已开源。
原文摘要 · Abstract (English)
Many classic opera videos exhibit poor visual quality due to the limitations of early filming equipment and long-term degradation during storage. Although real-world video super-resolution (RWVSR) has achieved significant advances in recent years, directly applying existing methods to degraded opera videos remains challenging. The difficulties are twofold. First, accurately modeling real-world degradations is complex: simplistic combinations of classical degradation kernels fail to capture the authentic noise distribution, while methods that extract real noise patches from external datasets are prone to style mismatches that introduce visual artifacts. Second, current RWVSR methods, which rely solely on degraded image features, struggle to reconstruct realistic and detailed textures due to a lack of high-level semantic guidance. To address these issues, we propose a Text-guided Dual-Branch Opera Video Super-Resolution (TextOVSR) network, which introduces two types of textual prompts to guide the super-resolution process. Specifically, degradation-descriptive text, derived from the degradation process, is incorporated into the negative branch to constrain the solution space. Simultaneously, content-descriptive text is incorporated into a positive branch and our proposed Text-Enhanced Discriminator (TED) to provide semantic guidance for enhanced texture reconstruction. Furthermore, we design a Degradation-Robust Feature Fusion (DRF) module to facilitate cross-modal feature fusion while suppressing degradation interference. Experiments on our OperaLQ benchmark show that TextOVSR outperforms state-of-the-art methods both qualitatively and quantitatively. The code is available at https://github.com/ChangHua0/TextOVSR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。