优化相位掩模可提升受限读出下的图像分类性能。
End-to-End Optimization of Incoherent Imaging for Classification Under Detector-Limited Readout

- 联合优化光学相位掩模与神经网络,提升分类效果。
- 在检测器读出受限时,性能显著优于传统透镜。
- 适合低噪声、低频特征主导的分类任务。
端到端联合优化光学前端(如超表面)与神经网络后端已广泛应用于成像任务,但缺乏对这类系统何时及为何优于传统透镜成像的理论框架。本文聚焦物体分类这一核心任务,研究在相干成像中,相位掩模的端到端优化何时能超越传统聚焦透镜。研究发现,性能增益主要出现在检测器读出受限的情况下;而在完全读出条件下,任何非相干相位掩模的性能均无法超过理想信道的互信息上限,传统透镜已接近该上限,联合优化无实际收益。当检测器读出受限(如空间采样粗糙或测量次数有限)时,优化光学元件可通过增强检测器输出中的类别可分性,显著提升分类性能。该增益在低噪声条件下最大,随噪声增加而减弱,因光学仅影响信号在进入检测器前的分布,无法消除后续噪声。此外,增益还依赖于任务的频谱结构:当类间差异集中在低频而类内变化在高频时,协同设计效果最佳。本文建立了理论框架,并在合成数据及标准基准(MNIST、FashionMNIST、SVHN)上验证了其预测能力。
原文摘要 · Abstract (English)
End-to-end co-optimization of optical front-ends (e.g. metasurfaces) and neural network back-ends has been widely applied to imaging tasks, yet a formalism characterizing when and why such systems outperform conventional lens-based imaging is largely lacking. This paper focuses on object classification, a central imaging task, and asks when end-to-end optimization of a phase mask for incoherent imaging improves performance over a conventional focusing lens. We find that these gains arise primarily under constrained detector readout and are limited under full detector readout. In the latter setting, we prove that no incoherent phase mask exceeds the ideal-channel mutual information between detector measurements and class labels; a conventional focusing lens approaches this ceiling, and joint optimization yields no empirical gain. When detector readout is constrained -- by coarse spatial sampling or a limited number of measurements -- optimized optics can substantially improve classification by increasing class separability in the detector measurements. These gains are largest under low detector noise and shrink as noise grows, because the optics shape the signal before it reaches the detector but cannot remove noise added afterward. The advantage also depends on the spectral structure of the task: co-design helps most when class-discriminative content is concentrated at lower spatial frequencies than within-class variation. We develop a theoretical framework formalizing these distinctions and test its predictions on synthetic data and standard benchmarks (MNIST, FashionMNIST, SVHN).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。