arXiv:2606.26738cs.CV2026-06

测试主流图像编辑模型对真实光照的理解能力,发现其仍有明显不足。

Do Image Editing Models Understand Lighting?

论文配图:Do Image Editing Models Understand Lighting?
图 1 · 摘自论文原文
  • 构建真实世界光照数据集3DLP,包含1000组室内场景的光探头开闭图像对
  • 模型在暗部区域的光照编辑误差显著高于亮部,金属表面更易出错
  • 现有视觉语言模型不适用于像素级光照传输分析,需专用评估方法

尽管生成式图像编辑模型在视觉保真度上取得进展,但其是否具备真实光照的内在理解仍未知。现有基准多依赖人工判断或合成数据,难以准确评估光照物理一致性。本文提出3D-anchored Light Probe(3DLP)基准,构建了包含1000组真实室内场景的高保真HDR数据集,其中光探头被实际开启与关闭。为支持细粒度分析,我们标注了投射阴影、金属表面等特定区域。通过新设计的两个评分指标,消除白平衡等生成伪影影响,评估多种前沿图像编辑模型的光照编辑准确性。结果显示,模型整体表现差异显著,镜面高光误差较小;最佳模型虽接近真实物理规律,但仍存改进空间。所有模型在接收光探头照射较少的区域均更易出错。此外,测试视觉语言模型在该任务上的表现,发现其不适合像素级光照传输分析。我们将基准与数据集公开,以推动相关研究。

原文摘要 · Abstract (English)

While recent advancements in generative image editing models have achieved stunning visual fidelity, it remains an open question whether these systems possess an intrinsic knowledge of real-world lighting. Existing benchmarks typically evaluate high-level plausibility of perceptual light transport on curated internet imagery, using VLMs or human judgement, or they rely on synthetically generated datasets. In this work, we introduce the 3D-anchored Light Probe (3DLP) benchmark, for which we have captured a new high-fidelity HDR dataset of real-world lighting changes. The dataset consists of 1K image pairs of diverse indoor scenery in which light probes are physically turned on and off. To allow for a granular performance analysis, we annotated specific image regions such as cast shadows or metallic surfaces. With this data, we evaluate a range of state-of-the-art image editing models by measuring how well their light probe edits align with reality. The evaluation uses two new scores to compensate for AI-generated photographic effects, such as adjusted white balance. Our results show that the overall performance of models differs considerably, with differences slightly less pronounced for specular highlights. The best image editing models are remarkably consistent with real-world physics, however, they still leave room for improvement. We observe that image regions that receive less light from the light probe are more prone to errors for all models. Furthermore, building on their success in evaluating macroscopic lighting plausibility, we test VLMs on our task but find that they are unsuitable for pixel-level light transport analysis. We will make the benchmark, together with the real-world dataset, publicly available to encourage future research on this topic.

光照理解图像编辑真实感生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。