用可见性约束提升阴影生成真实感,解决复杂场景中阴影与物体不一致问题。
VSDiffusion: Taming Ill-Posed Shadow Generation via Visibility-Constrained Diffusion
- 分两阶段生成:先粗略定位阴影区域,再结合光照深度信息精细化生成
- 在DESOBAv2数据集上多项指标达到新最优,阴影几何一致性显著提升
- 通过可见性控制和高频增强模块,精准处理边界与背景纹理交互
在图像合成中为插入的前景物体生成逼真投影阴影是一项关键但具挑战性的任务,由于阴影形成过程本质上的病态性,保持阴影与物体在复杂场景中的几何一致性仍很困难。为此,我们提出VSDiffusion,一种基于可见性约束的两阶段框架,通过引入可见性先验来缩小解空间。第一阶段预测粗略阴影掩码以定位可能生成阴影的区域;第二阶段则在复合图像的光照和深度线索引导下进行条件扩散生成精确阴影。VSDiffusion通过两条互补路径注入可见性先验:一是带有阴影门控交叉注意力的可见性控制分支,提供多尺度结构引导;二是学习得到的软先验图,在易错区域重加权训练损失以增强几何修正能力。此外,还引入高频引导增强模块,以锐化边界并改善与背景的纹理交互。在广泛使用的公共DESOBAv2数据集上的实验表明,所提出的VSDiffusion能生成准确的阴影,并在多数评估指标上达到新的最先进水平。
原文摘要 · Abstract (English)
Generating realistic cast shadows for inserted foreground objects is a crucial yet challenging problem in image composition, where maintaining geometric consistency of shadow and object in complex scenes remains difficult due to the ill-posed nature of shadow formation. To address this issue, we propose VSDiffusion, a visibility-constrained two-stage framework designed to narrow the solution space by incorporating visibility priors. In Stage I, we predict a coarse shadow mask to localize plausible shadow generated regions. And in Stage II, conditional diffusion is performed guided by lighting and depth cues estimated from the composite to generate accurate shadows. In VSDiffusion, we inject visibility priors through two complementary pathways. First, a visibility control branch with shadow-gated cross attention that provides multi-scale structural guidance. Then, a learned soft prior map that reweights training loss in error-prone regions to enhance geometric correction. Additionally, we also introduce high-frequency guided enhancement module to sharpen boundaries and improve texture interaction with the background. Experiments on widely used public DESOBAv2 dataset demonstrated that our proposed VSDiffusion can generate accurate shadow, and establishes new SOTA results across most evaluation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。