arXiv:2606.19985cs.CV2026-06

用视觉语言模型修复光场图像中的遮挡,提升复杂场景可见性。

Vision-Reasoning-Guided Occlusion Removal from Light Fields

论文配图:Vision-Reasoning-Guided Occlusion Removal from Light Fields
图 1 · 摘自论文原文
  • 结合光场整合与视觉语言模型,利用语义先验指导遮挡恢复。
  • 在4个合成场景上平均SSIM达当前最优,真实数据集表现稳定。
  • 适合搜救、机器人导航等需高鲁棒性视觉的场景应用。

遮挡鲁棒的场景重建仍是计算成像中的重大挑战,尤其在密集植被严重限制可视性的场景中。本文提出一种视觉-推理引导的光场遮挡去除框架,将光场积分(LFI)与视觉语言模型(VLM)语义推理相结合。首先通过LFI融合多视角观测,抑制前景遮挡,生成初步增强可见性的表征;随后,VLM作为条件语义先验,恢复退化的结构与细节。采用多样本融合策略聚合多个生成假设,提升结果一致性并减少幻觉。在合成与真实世界数据集上的实验表明,该方法在四个合成基准场景(4-Syn)上实现最高平均SSIM,且在结构化与非结构化采集设置下均表现出强泛化能力,适用于搜救及探索型机器人导航任务。

原文摘要 · Abstract (English)

Occlusion-robust scene recovery remains a major challenge in computational imaging, particularly where dense vegetation severely limits visibility. We propose a visionreasoning-guided light field occlusion removal framework combining light field integration (LFI) with vision-language model (VLM) semantic reasoning. Multi-view observations are first integrated via LFI to suppress foreground occlusions, producing an initial visibility-enhanced representation, a VLM then acts as a conditional semantic prior to restore degraded structures and fine details. A multi-sample fusion strategy aggregates multiple generated hypotheses to improve consistency and reduce hallucination. Experimental results on synthetic and real-world datasets show state-of-the-art performance, achieving the highest average SSIM across four synthetic benchmark scenes (4-Syn) and strong generalization across structured and unstructured acquisition settings, with applicability to search-and-rescue and exploratory robotic navigation.

光场成像遮挡恢复视觉语言模型机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。