系统梳理深度学习时代阴影检测、去除与生成的现状与挑战
Unveiling Deep Shadows: A Survey and Benchmark on Image and Video Shadow Detection, Removal, and Generation in the Deep Learning Era
- 构建统一分类体系与标准化评估基准
- 发现模型性能受分辨率和数据集偏差显著影响
- 适合研究视觉感知、图像修复与AIGC真实性的学者
阴影由光线遮挡形成,在视觉感知中起关键作用,直接影响场景理解、图像质量与视觉真实感。本文首次系统综述并构建了图像与视频领域基于深度学习的阴影检测、去除与生成的统一调研与基准测试。提出一致的架构、监督策略与学习范式分类体系;梳理主流数据集与评估协议;在标准化设置下重训练代表性方法以实现公平比较。基准结果揭示:先前研究存在报告不一致、模型设计与分辨率依赖性强、跨数据集泛化能力有限等问题。通过整合三类任务的洞察,强调光照共性线索与先验的连接作用。展望未来方向包括一体化统一框架、语义与几何感知推理、基于阴影的AIGC真实性分析,以及物理引导先验融入多模态基础模型。已公开修正后的数据集、训练模型与评估工具,支持可复现研究。
原文摘要 · Abstract (English)
Shadows, formed by the occlusion of light, play an essential role in visual perception and directly influence scene understanding, image quality, and visual realism. This paper presents a unified survey and benchmark of deep-learning-based shadow detection, removal, and generation across images and videos. We introduce consistent taxonomies for architectures, supervision strategies, and learning paradigms; review major datasets and evaluation protocols; and re-train representative methods under standardized settings to enable fair comparison. Our benchmark reveals key findings, including inconsistencies in prior reports, strong dependence on model design and resolution, and limited cross-dataset generalization due to dataset bias. By synthesizing insights across the three tasks, we highlight shared illumination cues and priors that connect detection, removal, and generation. We further outline future directions involving unified all-in-one frameworks, semantics- and geometry-aware reasoning, shadow-based AIGC authenticity analysis, and the integration of physics-guided priors into multimodal foundation models. Corrected datasets, trained models, and evaluation tools are released to support reproducible research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。