arXiv:2512.15701cs.CV2025-12被引 1

用视觉语言模型做图像压缩的感知评判,零样本效果媲美人类。

VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression

  • 用VLM直接判断图像对差异,零样本替代传统感知损失
  • 在多个数据集上达到或超过当前最佳的人类感知压缩效果
  • 无需额外训练感知网络,适合追求真实人类偏好评估的研究者

图像压缩的评价若包含人类偏好,通常发现均方误差等简单失真函数与人类感知不一致。为使压缩模型更符合人类感知,以往工作采用基于大规模人类心理视觉判断数据集校准的可微分感知损失。本文发现,令人惊讶的是,最先进的视觉语言模型(VLM)在未经过任何微调的情况下,能零样本准确复现人类二选一强制选择(2AFC)判断。受此启发,我们提出视觉语言模型用于图像压缩(VLIC),一种基于扩散模型的压缩系统,通过二元VLM判断进行后训练。该方法利用现有扩散模型偏好后训练技术,而非将VLM判断蒸馏为独立感知损失网络。实验表明,在多种数据集上,基于VLM判断校准的VLIC在感知指标和大规模用户研究中表现优异,达到竞争性或领先水平。我们还对基于VLM的奖励设计与训练流程进行了深入分析,并分享了关键洞察。更多可视化内容见 https://kylesargent.github.io/vlic

原文摘要 · Abstract (English)

Evaluations of image compression performance which include human preferences have generally found that naive distortion functions such as MSE are insufficiently aligned to human perception. In order to align compression models to human perception, prior work has employed differentiable perceptual losses consisting of neural networks calibrated on large-scale datasets of human psycho-visual judgments. We show that, surprisingly, state-of-the-art vision-language models (VLMs) can replicate binary human two-alternative forced choice (2AFC) judgments zero-shot when asked to reason about the differences between pairs of images. Motivated to exploit the powerful zero-shot visual reasoning capabilities of VLMs, we propose Vision-Language Models for Image Compression (VLIC), a diffusion-based image compression system designed to be post-trained with binary VLM judgments. VLIC leverages existing techniques for diffusion model post-training with preferences, rather than distilling the VLM judgments into a separate perceptual loss network. We show that calibrating this system on VLM judgments produces competitive or state-of-the-art performance on human-aligned visual compression depending on the dataset, according to perceptual metrics and large-scale user studies. We additionally conduct an extensive analysis of the VLM-based reward design and training procedure and share important insights. More visuals are available at https://kylesargent.github.io/vlic

图像压缩视觉语言模型感知评估扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。