MILO是轻量级图像质量评估模型,可实时优化生成图像的视觉效果。
MILO: A Lightweight Perceptual Quality Metric for Image and Latent-Space Optimization
- 用伪MOS数据训练,无需人工标注,结合多尺度感知机制。
- 在标准评测中超越现有指标,推理速度快,适合实时应用。
- 可作为感知损失用于图像与隐空间优化,提升修复和超分效果。
我们提出MILO(图像与隐空间优化度量),一种轻量级、多尺度的全参考图像质量评估(FR-IQA)方法。MILO通过伪均值意见分数(MOS)监督训练,对多样化图像施加可复现的失真,并利用近期质量评估指标的集成评分,考虑了视觉掩蔽效应。该方法避免了大规模人工标注数据需求。尽管架构紧凑,MILO在标准FR-IQA基准上表现优于现有指标,且推理速度快,适用于实时场景。除了质量预测,我们还展示了MILO作为图像与隐空间感知损失的实用性。特别地,在Stable Diffusion的VAE编码器输出的隐表示中引入空间掩蔽建模,实现了高效且视觉一致的优化。结合课程学习策略,先处理感知不敏感区域,再逐步聚焦于视觉失真更显著区域,显著提升了去噪、超分辨率和人脸修复等任务性能,同时降低计算开销。因此,MILO既是先进图像质量度量工具,也是生成流水线中实用的感知优化手段。
原文摘要 · Abstract (English)
We present MILO (Metric for Image- and Latent-space Optimization), a lightweight, multiscale, perceptual metric for full-reference image quality assessment (FR-IQA). MILO is trained using pseudo-MOS (Mean Opinion Score) supervision, in which reproducible distortions are applied to diverse images and scored via an ensemble of recent quality metrics that account for visual masking effects. This approach enables accurate learning without requiring large-scale human-labeled datasets. Despite its compact architecture, MILO outperforms existing metrics across standard FR-IQA benchmarks and offers fast inference suitable for real-time applications. Beyond quality prediction, we demonstrate the utility of MILO as a perceptual loss in both image and latent domains. In particular, we show that spatial masking modeled by MILO, when applied to latent representations from a VAE encoder within Stable Diffusion, enables efficient and perceptually aligned optimization. By combining spatial masking with a curriculum learning strategy, we first process perceptually less relevant regions before progressively shifting the optimization to more visually distorted areas. This strategy leads to significantly improved performance in tasks like denoising, super-resolution, and face restoration, while also reducing computational overhead. MILO thus functions as both a state-of-the-art image quality metric and as a practical tool for perceptual optimization in generative pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。