让镜头设计贴合冻结的视觉模型,提升小尺寸光学系统的识别准确率。
VLM-Aware Meta-Optic Front-End Design for Frozen Vision-Language Models

- 用可微成像与麦克斯韦仿真联合优化连续密度超表面镜头。
- 在ImageNet-100上将CLIP零样本准确率从53.75%提升至65.41%。
- 成果可跨模型迁移,适合微型化智能感知系统设计者参考。
传统机器视觉管道依赖高质量光学器件以生成清晰、人可理解的图像,光学设计长期基于分辨率、像差校正和像素保真度等图像级标准。然而,在尺寸、成本或外形受限的应用中,这些光学器件往往不切实际,紧凑型超表面光学成为替代方案,但受物理效率严格限制。本文提出CODA框架,通过可微成像与基于麦克斯韦方程的伴随梯度更新,联合优化连续密度超表面前端,直接针对固定零样本CLIP分类器的交叉熵损失进行优化,无需学习重建、图像信号处理或图像保真辅助目标。在ImageNet-100的二维模拟成像基准上,相比焦聚基线,CODA将CLIP ViT-L/14的零样本准确率从53.75±3.57%提升至65.41±3.99%。优化后的光学器件在未重新优化的情况下,可跨模型迁移至CLIP、SigLIP和DINOv2,在ImageNet-100、CIFAR-100和Food-101上均表现良好。结果表明,在受限超表面成像条件下,将光学设计与冻结视觉模型目标对齐,可显著提升下游识别性能,而非依赖传统图像形成标准。
原文摘要 · Abstract (English)
Conventional machine-vision pipelines typically rely on high-quality optics that produce clean, human-interpretable images, and optical design has therefore been driven by image-level criteria such as resolution, aberration correction, and pixel fidelity. However, such optics are often impractical for size-, cost-, or form-factor-constrained applications, where compact meta-optics offer an attractive alternative but operate under strict physical efficiency limits. We propose CODA, a co-design framework that optimizes a continuous-density meta-optic front-end for frozen-model recognition using differentiable image formation and adjoint-gradient updates of Maxwell-based simulations. CODA directly optimizes the cross-entropy loss of a fixed zero-shot CLIP classifier without learned reconstruction, image signal processing, or image-fidelity auxiliary objectives. In a two-dimensional simulated imaging benchmark on ImageNet-100, CODA improves CLIP ViT-L/14 zero-shot accuracy from 53.75 $\pm$ 3.57$\%$ with a focal-concentration baseline to 65.41 $\pm$ 3.99$\%$. The optimized optics further transfer without re-optimization across CLIP, SigLIP, and DINOv2 on ImageNet-100, CIFAR-100, and Food-101. These results demonstrate that, under constrained meta-optic imaging, downstream recognition can be improved by aligning optical design with frozen vision-model objectives rather than conventional image-formation criteria.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。