用真人评价优化超薄透镜成像质量评估,大幅减少标注量。
MetaRanker: Human-in-the-loop Active Ranking for Metalens Image Quality

- 结合人类偏好与视觉语言模型,主动选择有信息量的图像对比
- 在真实和合成数据上,排名与人眼判断高度一致,标注量降80%
- 针对金属透镜成像特性,以可识别性为核心评价指标
现代成像系统中图像质量由传感器、光学与计算重建共同决定。超薄金属透镜虽能显著缩小光学模组体积,但实际设计常存在明显的色差与视场依赖像差,需计算重建。当前金属透镜流程多采用基于失真度的指标(如PSNR)训练与筛选重建模型,但此类指标与人类偏好及下游使用价值相关性弱,存在感知-失真权衡问题。本文提出MetaRanker,一种人机协同的主动排序框架,将金属透镜图像质量定义为语义可解释性——即在光学伪影存在下,人类可靠识别物体与结构的能力。MetaRanker融合概率偏好模型与不确定性感知查询选择,并利用视觉-语言模型提供轻量级语义先验,仅用于引导信息丰富对比的采样,人类判断始终为唯一监督信号。在具有不同退化特征的真实与合成金属透镜数据集上,MetaRanker生成的排名与人类评估最一致,且相比全量成对标注,所需标注数减少约80%。此外,我们发现标准图像质量评估指标在金属透镜领域与人类可解释性关联有限,表明MetaRanker为感知驱动的金属透镜评估与协同设计提供了实用路径。
原文摘要 · Abstract (English)
Image quality in modern imaging systems emerges from the coupled effects of the sensor, optics, and computational reconstruction. Ultra-thin metalenses offer a path toward substantial miniaturization of optical modules, but practical designs often exhibit pronounced chromatic and field-dependent aberrations that necessitate computational reconstruction. In current metalens pipelines, reconstruction models are commonly trained and selected using distortion-based fidelity objectives, such as PSNR, yet these proxies can be weakly correlated with human preference and downstream utility, reflecting the well-known perception--distortion trade-off. We introduce MetaRanker, a human-in-the-loop active ranking framework that formalizes metalens image quality in terms of semantic interpretability, defined as the degree to which humans can reliably recognize objects and structures in the presence of optical artifacts. MetaRanker combines a probabilistic preference model with uncertainty-aware query selection, and leverages vision--language models to provide lightweight semantic priors. Importantly, these priors are used only to guide the sampling of informative comparisons; human judgments remain the primary supervision signal throughout. Across real-world and synthetic metalens datasets with distinct degradation profiles, MetaRanker produces rankings that align most closely with human assessments, while reducing the number of pairwise annotations required by approximately 80% relative to exhaustive pairwise evaluation. Finally, we show that standard image quality assessment metrics exhibit limited alignment with human interpretability in the metalens domain, positioning MetaRanker as a practical step toward perceptually grounded metalens evaluation and co-design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。