arXiv:2602.18720cs.CV2026-02

解决静态图像中细微运动模糊的检测与分割难题,提升视频封面画质评估能力。

Subtle Motion Blur Detection and Segmentation from Static Image Artworks

  • 基于SAM分割区域模拟可控相机与物体运动,生成真实细微模糊图像及精确标注
  • 零样本检测在GoPro上准确率达89.68%,在CUHK上平均交并比达59.77%(基线9.00%)
  • 适合视频平台自动质检、智能裁剪和画质过滤场景

流媒体服务全球覆盖数亿用户,缩略图、海报图等视觉资产对用户点击至关重要。细微运动模糊普遍存在,降低画面清晰度,损害用户信任与点击率。然而,从静态图像中检测运动模糊仍缺乏研究,现有方法与数据集多关注严重模糊,且缺少像素级精细标注。如GOPRO与NFS等基准常含强合成模糊,其清晰参考图中仍有残余模糊,导致监督信号模糊。我们提出SMBlurDetect,一个统一框架,结合高质量运动模糊专用数据集生成与端到端检测器,支持多粒度零样本检测。该流程利用超高清美学图像,通过控制相机与物体运动仿真,在SAM分割区域上合成真实细微模糊,结合透明度感知合成与均衡采样,生成空间局部化、精准标注的模糊图像。采用基于U-Net的检测器,使用ImageNet预训练编码器,融合课程学习、难样本挖掘、焦点损失、模糊频率通道与分辨率感知增强的混合策略进行训练。本方法实现强大零样本泛化:在GoPro上准确率达89.68%(基线66.50%),在CUHK上平均交并比达59.77%(基线9.00%),分割性能提升6.6倍。定性结果表明能准确定位细微模糊伪影,可用于低质量帧自动过滤与智能裁剪中的兴趣区域提取。

原文摘要 · Abstract (English)

Streaming services serve hundreds of millions of viewers worldwide, where visual assets such as thumbnails, box art, and cover images are critical for engagement. Subtle motion blur remains a pervasive quality issue, reducing visual clarity and negatively affecting user trust and click-through rates. However, motion blur detection from static images is underexplored, as existing methods and datasets focus on severe blur and lack fine-grained pixel-level annotations needed for quality-critical applications. Benchmarks such as GOPRO and NFS are dominated by strong synthetic blur and often contain residual blur in their sharp references, leading to ambiguous supervision. We propose SMBlurDetect, a unified framework combining high-quality motion blur specific dataset generation with an end-to-end detector capable of zero-shot detection at multiple granularities. Our pipeline synthesizes realistic motion blur from super high resolution aesthetic images using controllable camera and object motion simulations over SAM segmented regions, enhanced with alpha-aware compositing and balanced sampling to generate subtle, spatially localized blur with precise ground truth masks. We train a U-Net based detector with ImageNet pretrained encoders using a hybrid mask and image centric strategy incorporating curriculum learning, hard negatives, focal loss, blur frequency channels, and resolution aware augmentation.Our method achieves strong zero-shot generalization, reaching 89.68% accuracy on GoPro (vs 66.50% baseline) and 59.77% Mean IoU on CUHK (vs 9.00% baseline), demonstrating 6.6x improvement in segmentation. Qualitative results show accurate localization of subtle blur artifacts, enabling automated filtering of low quality frames and precise region of interest extraction for intelligent cropping.

运动模糊图像质量零样本智能裁剪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。