根据内容动态选配元数据,让生成式超分更高效清晰。
MetaSR: Content-Adaptive Metadata Orchestration for Generative Super-Resolution

- 用DiT模型自适应筛选并注入相关元数据
- 在相同画质下节省50%传输码率,最高提升1.0dB PSNR
- 适合视频/图像超分中内容多变、带宽受限的场景
我们研究真实场景下的生成式超分辨率(SR),其中内容与退化类型在不同领域、题材和片段间差异显著。例如,图像与视频可能交替出现文字叠加、快速运动、平滑动画和低光照人脸,每类均需不同的辅助信息。现有元数据引导的SR方法通常采用固定条件设计,在有用线索依赖内容且传输预算有限时表现不佳。我们提出MetaSR,一种基于扩散变换器(DiT)的框架,能在资源受限下选择并注入任务相关的元数据以指导超分。具体地,利用DiT自身的VAE和变压器主干融合异构元数据,并采用高效蒸馏策略实现一步扩散推理。在多种内容类别和退化模式下的实验表明,MetaSR相比基准方案在相同质量下可节省高达50%的传输码率,同时提升1.0~1.0 dB PSNR。我们在率-失真优化(RDO)框架下评估这些收益,联合考虑发送端码率与接收端/显示端质量指标(如PSNR和SSIM)。
原文摘要 · Abstract (English)
We study generative super-resolution (SR) in real-world scenarios where content and degradations vary across domains, genres, and segments. For example, images and videos may alternate between text overlays, fast motion, smooth cartoons, and low-light faces, each benefiting from different forms of side information. Existing metadata-guided SR methods typically use a fixed conditioning design, which is suboptimal when useful cues are content dependent and transmission budgets are limited. We propose MetaSR, a Diffusion Transformer (DiT)-based framework that selects and injects task-relevant metadata to guide SR under resource constraints. Specifically, we use the DiT's own VAE and transformer backbone to fuse heterogeneous metadata, and adopt an efficient distillation strategy that enables one-step diffusion inference. Experiments across diverse content buckets and degradation regimes show that MetaSR outperforms reference solutions by up to 1.0~dB PSNR while achieving up to 50\% transmission bitrate saving at matched quality. We assess these gains under a rate--distortion optimization (RDO) framework that jointly accounts for sender-side bitrate and receiver/display quality metrics (e.g., PSNR and SSIM).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。