arXiv:2601.22189eess.IVcs.CV2026-01中稿 · ICASSP 2026

用语义信息优化视频压缩,让重要细节更清晰。

SCENE: Semantic-aware Codec Enhancement with Neural Embeddings

  • 将视觉语言模型的语义嵌入融入轻量卷积网络,优先保护关键区域。
  • 在高分辨率数据集上,MS-SSIM和VMAF指标均优于基线方法。
  • 无需改动现有编码流程,可实时运行,适合部署到真实系统。

标准视频编码器产生的压缩伪影会降低主观质量。本文提出一种轻量级、语义感知的预处理框架,通过选择性修复这些失真来提升感知保真度。方法将视觉语言模型生成的语义嵌入整合进高效的卷积架构中,优先保留具有感知重要性的结构。模型采用可微分的编码器代理进行端到端训练,能有效缓解多种标准编码器的伪影,且无需修改现有视频流水线。推理时丢弃编码器代理,SCENE作为独立预处理器运行,实现实时性能。在高分辨率基准测试中,该方法在客观(MS-SSIM)和主观(VMAF)指标上均优于基线,尤其在显著区域的细节纹理保留上表现突出。结果表明,语义引导的、编码器感知的预处理是增强压缩视频流的有效方法。

原文摘要 · Abstract (English)

Compression artifacts from standard video codecs often degrade perceptual quality. We propose a lightweight, semantic-aware pre-processing framework that enhances perceptual fidelity by selectively addressing these distortions. Our method integrates semantic embeddings from a vision-language model into an efficient convolutional architecture, prioritizing the preservation of perceptually significant structures. The model is trained end-to-end with a differentiable codec proxy, enabling it to mitigate artifacts from various standard codecs without modifying the existing video pipeline. During inference, the codec proxy is discarded, and SCENE operates as a standalone pre-processor, enabling real-time performance. Experiments on high-resolution benchmarks show improved performance over baselines in both objective (MS-SSIM) and perceptual (VMAF) metrics, with notable gains in preserving detailed textures within salient regions. Our results show that semantic-guided, codec-aware pre-processing is an effective approach for enhancing compressed video streams.

视频压缩语义感知感知质量轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。