arXiv:2512.06174cs.CV2025-12

用物理规则生成更真实的阴影,让光影关系更合理。

Embedding Physical Reasoning into Diffusion-Based Shadow Generation

  • 基于几何推理恢复场景结构和光照方向,提供阴影位置依据。
  • 在DESOBAV2数据集上,阴影定位误差降低23%,误判率降30%。
  • 适合需要高精度阴影合成的三维重建与图形渲染任务。

为插入物体生成真实阴影需理解场景几何与光照关系。但现有方法多在图像空间操作,隐式学习物体、光照与阴影间的物理联系,常导致阴影错位或不自然。本文将阴影生成建立在阴影形成的物理原理之上:给定合成图像和物体掩码,通过恢复近似场景几何并估计主光源方向,利用几何推理生成物理一致的阴影初值。该初值虽粗略,但提供了阴影位置的参考。由于单图无法唯一确定光照,我们预测光照与阴影线索的置信度,并据此调节其在生成中的影响。这些线索——阴影掩码、光照方向及其置信度——作为条件输入扩散模型,将其精炼为逼真阴影。在DESOBAV2上的实验表明,本方法在阴影真实感与定位准确性上均优于当前最优,阴影区域均方根误差降低23%,误判率降低30%。

原文摘要 · Abstract (English)

Generating realistic shadows for inserted objects requires reasoning about scene geometry and illumination. However, most existing methods operate purely in image space, leaving the physical relationship between objects, lighting, and shadows to be learned implicitly, often resulting in misaligned or implausible shadows. We instead ground shadow generation in the physics of shadow formation. Given a composite image and an object mask, we recover approximate scene geometry and estimate a dominant light direction to derive a physics-grounded shadow estimate via geometric reasoning. While coarse, this estimate provides a spatial anchor for shadow placement. Because illumination cannot always be uniquely inferred from a single image, we predict confidence scores for both lighting and shadow cues and use them to regulate their influence during generation. These cues, shadow mask, light direction, and their confidences, condition a diffusion-based generator that refines the estimate into a realistic shadow. Experiments on DESOBAV2 show that our method improves both shadow realism and localization, achieving 23% lower shadow-region RMSE and 30% lower shadow-region BER over prior state-of-the-art.

阴影生成扩散模型物理建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。