arXiv:2503.11181cs.CVcs.AI2025-03被引 1

用扩散模型分阶段提升足球直播画质,64×64变1024×1024。

Multi-Stage Generative Upscaler: Reconstructing Football Broadcast Images via Diffusion Models

  • 分三阶段生成,结合ControlNet和定制LoRA控制细节。
  • 在64×64输入下实现1024×1024高清输出,还原球衣标志等细节。
  • 专为体育广播优化,适合视频增强与实时分析场景。

低分辨率足球直播图像的重建在体育转播中面临巨大挑战,因细节视觉对分析与观众体验至关重要。本文提出一种基于扩散模型的多阶段生成超分辨率框架,将最低达64×64像素的输入图像重构为1024×1024的高保真输出。通过整合图像到图像管道、ControlNet条件控制及LoRA微调,该方法显著优于传统超分技术,在恢复复杂纹理与特定领域元素(如球员细节、球衣徽标)方面表现突出。定制的LoRA在自建足球数据集上训练,确保对体育广播需求的适配性。实验表明,ControlNet有效提升细节精度,LoRA增强任务相关特征。结果验证了扩散模型在体育媒体图像重建中的潜力,为自动视频增强与实时体育分析提供了新路径。

原文摘要 · Abstract (English)

The reconstruction of low-resolution football broadcast images presents a significant challenge in sports broadcasting, where detailed visuals are essential for analysis and audience engagement. This study introduces a multi-stage generative upscaling framework leveraging Diffusion Models to enhance degraded images, transforming inputs as small as $64 \times 64$ pixels into high-fidelity $1024 \times 1024$ outputs. By integrating an image-to-image pipeline, ControlNet conditioning, and LoRA fine-tuning, our approach surpasses traditional upscaling methods in restoring intricate textures and domain-specific elements such as player details and jersey logos. The custom LoRA is trained on a custom football dataset, ensuring adaptability to sports broadcast needs. Experimental results demonstrate substantial improvements over conventional models, with ControlNet refining fine details and LoRA enhancing task-specific elements. These findings highlight the potential of diffusion-based image reconstruction in sports media, paving the way for future applications in automated video enhancement and real-time sports analytics.

图像修复扩散模型体育视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。