arXiv:2503.00747cs.CVcs.RO2025-03

首个统一光场语义分割与显著目标检测框架,适配多种光场表示。

LFX: Towards Unified Light Field Dense Semantic Segmentation and Salient Object Detection

  • 构建不变表示的特征调制空间,支持多类光场输入
  • 在3个基准上实现最高性能,显著目标检测MAE低至0.027
  • 适合研究光场感知与跨表示学习的开发者

光场相机可在单次曝光中捕捉多视角观测。然而,现有研究通常针对特定光场表示设计,缺乏统一的学习框架。为此,我们提出LFX,首个面向光场感知的统一框架。LFX建立表示无关的特征调制空间,使其可适应异构光场表示及多样感知任务。具体地,我们提出视差角子空间建模(FoP-ASM),为每个辅助视图分配独立的角标记,实现视图级独立建模;同时,共享流形子空间约束与正则化损失确保跨视图全局一致的语义调制。在三个光场基准上的广泛评估表明,LFX在不同光场表示下均达到领先性能,相比特定表示方法提升高达12%和20%,显著目标检测的MAE分别降至0.029/0.027,语义分割达到84.37 mIoU。源代码将公开于https://github.com/FeiT-FeiTeng/LFX。

原文摘要 · Abstract (English)

Light field cameras capture multi-view observations within a single exposure. However, existing studies are typically tailored to specific LF representations, leaving the field without a unified learning framework. To bridge this gap, we present LFX, the first unified framework for LF perception. LFX establishes a representation-invariant feature modulation space, enabling it to adapt to heterogeneous LF representations and diverse perception tasks. Specifically, we propose Field-of-Parallax Angular Subspace Modeling (FoP-ASM), which assigns an independent angular marker to each auxiliary view, enabling view-wise independent modeling. Meanwhile, shared manifold subspace constraints and regularization losses enforce globally consistent semantic modulation across views. Extensive evaluations across three LF benchmarks show that LFX achieves state-of-the-art results across distinct LF representations, outperforming representation-specific methods by up to 12% and 20% with 0.029/0.027 MAE for salient object detection, and achieving 84.37 mIoU for semantic segmentation. The source code will be made publicly available at https://github.com/FeiT-FeiTeng/LFX.

光场感知语义分割显著目标检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。