arXiv:2506.12269eess.IVcs.CV2025-06被引 6

针对视频会议的超分辨率挑战,提升低清视频画质。

ICME 2025 Grand Challenge on Video Super-Resolution for Video Conferencing

  • 用因果模型实现低延迟视频超分,支持多类视频内容。
  • 在三种视频类型上均显著提升主观画质,符合实时通信需求。
  • 开源屏幕内容数据集,推动该领域研究发展。

超分辨率(SR)是计算机视觉中的关键任务,旨在从低分辨率(LR)输入重建高分辨率(HR)图像。单图超分辨率已取得显著进展,而视频超分辨率(VSR)则扩展到时序域,通过局部、单向、双向传播或传统上采样结合修复等方法提升视频质量。本挑战聚焦于视频会议场景,处理以H.265编码、固定量化参数(QP)的低分辨率视频,目标是在特定放大倍数下生成高分辨率输出,并在低延迟条件下保持良好感知质量,采用因果模型。挑战包含三个赛道:通用视频、说话人头像视频和屏幕内容视频,主办方提供了训练、验证和测试专用数据集。本次挑战还公开了一个新的屏幕内容数据集,用于超分辨率任务。提交结果通过众包实现的ITU-T Rec P.910标准进行主观评测。

原文摘要 · Abstract (English)

Super-Resolution (SR) is a critical task in computer vision, focusing on reconstructing high-resolution (HR) images from low-resolution (LR) inputs. The field has seen significant progress through various challenges, particularly in single-image SR. Video Super-Resolution (VSR) extends this to the temporal domain, aiming to enhance video quality using methods like local, uni-, bi-directional propagation, or traditional upscaling followed by restoration. This challenge addresses VSR for conferencing, where LR videos are encoded with H.265 at fixed QPs. The goal is to upscale videos by a specific factor, providing HR outputs with enhanced perceptual quality under a low-delay scenario using causal models. The challenge included three tracks: general-purpose videos, talking head videos, and screen content videos, with separate datasets provided by the organizers for training, validation, and testing. We open-sourced a new screen content dataset for the SR task in this challenge. Submissions were evaluated through subjective tests using a crowdsourced implementation of the ITU-T Rec P.910.

视频超分视频会议因果模型数据集开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。