用傅里叶-局部混合结构直接预测光场视差,无需构建代价体。
Frequency-Structured Field Learning for Light-Field Disparity Estimation

- 通过傅里叶低频与局部卷积联合更新潜空间特征,实现全局一致与局部精细的平衡
- 在HCI基准上达到强监督基线精度,且基础模型不依赖显式代价体
- 适合追求高效高精度光场深度估计的研究者或工业应用
光场视差估计需在平滑或纹理缺失区域保持全局一致性,在遮挡边界、细结构和突变深度处实现局部精度。现有方法多依赖EPI匹配、代价体或焦点堆栈构造、视图聚合或直接卷积回归,常受限于局部窗口、离散视差假设、内存密集型体积或注意力聚合。本文提出在场层面进行视差估计,从全局与局部更新的EPI衍生潜在特征中直接预测视差,无需显式构建视差体。引入FreqLF,一种由EPI引导的傅里叶-局部框架,同时编码水平与垂直EPI堆栈的角偏移线索及中心视图外观特征。这些线索被投影至潜在场,并通过堆叠的混合傅里叶-局部层迭代更新:傅里叶低模态更新实现全局特征交互,局部卷积保留细节所需的空间变化。最后,坐标条件化的高斯混合解码器输出视差,以混合均值作为最终估计。在HCI 4D光场基准上的实验表明,FreqLF接近强监督基线的精度,同时基础模型避免了显式代价体构建。消融实验证实傅里叶与局部分支的互补作用,缩放实验展示其在不同空间分辨率下的良好表现。结果表明,傅里叶-局部潜场学习是光场视差估计的一种有竞争力的替代方案。代码将随后发布。
原文摘要 · Abstract (English)
Light-field disparity estimation requires global consistency in smooth or textureless regions and local precision near occlusion boundaries, thin structures, and abrupt depth transitions. Existing methods address these requirements through EPI matching, cost-volume or focal-stack construction, view aggregation, or direct convolutional regression, often relying on local windows, discrete disparity hypotheses, memory-intensive volumes, or attention-based aggregation. We instead formulate disparity estimation at the field level, predicting disparity from globally and locally updated EPI-derived latent features without explicitly constructing a disparity volume. We introduce FreqLF, an EPI-guided Fourier-local framework that encodes angular parallax cues from horizontal and vertical EPI stacks together with central-view appearance features. These cues are projected into a latent field and updated through stacked hybrid Fourier-local layers. Fourier low-mode updates enable global feature interaction, while local convolutions preserve spatial variations needed for fine disparity detail. A coordinate-conditioned Gaussian-mixture decoder then predicts disparity, using the mixture mean as the final estimate. Experiments on the HCI 4D Light Field Benchmark show that FreqLF approaches the accuracy of strong supervised baselines while avoiding explicit cost-volume construction in the base model. Ablations confirm the complementary roles of the Fourier and local branches, and scaling experiments demonstrate practical behavior across spatial resolutions. These results suggest that Fourier-local latent field learning is a competitive alternative for light-field disparity estimation. The code will be published soon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。