提出更真实的运动模糊检测基准与新模型,提升定位精度与跨数据集泛化能力。
BOCCHI: A More Realistic and Challenging Benchmark for Local Motion Blur Detection with MSDCT-UNet

- 设计真实拍摄的BOCCHI基准,避免梯度捷径依赖
- 新模型在本地和边界定位上均达最优,仅用633图训练
- 适合需要高精度模糊区域检测的研究者
局部运动模糊检测需精准定位模糊像素。现有基准使模型依赖梯度捷径,难以迁移。本文提出BOCCHI(Blurred Objects Captured across Cameras with Human-annotated Imagery),一个真实拍摄的基准,其清晰区域与模糊梯度分布重叠,有效阻断捷径。同时提出MSDCT-UNet(多尺度离散余弦变换UNet),通过DCT注意力与FiLM注入多尺度DCT先验。MSDCT-UNet在BOCCHI上实现域内mIoU与边界定位最优,且仅用633张图像训练的模型,在跨数据集迁移中超越所有其他训练源。
原文摘要 · Abstract (English)
Local motion blur detection requires pixel-level localization of blurred regions. Existing benchmarks let models rely on gradient shortcuts that fail to transfer. We introduce BOCCHI (Blurred Objects Captured across Cameras with Human-annotated Imagery), a real-captured benchmark whose sharp regions overlap the blur gradient distribution and defeat these shortcuts, and propose MSDCT-UNet (Multi-Scale Discrete Cosine Transform UNet), a frequency-aware encoder-decoder injecting multi-scale DCT priors through DCT Attention and FiLM. MSDCT-UNet ranks first in in-domain mIoU and boundary localization on BOCCHI, and BOCCHI-trained models outperform every other training source on cross-dataset transfer with only 633 training images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。