arXiv:2603.07272cs.AI2026-03

用图像质量变化自动学习视觉偏好,无需人工标注。

VisualDeltas: Learning Preferences from Visual Quality Perturbations

  • 通过图像质量扰动生成偏好信号,无需人工标注或外部教师。
  • 在多模态基准上优于拒绝采样微调,提升模型泛化能力。
  • 适用于多种视觉退化场景,灵活适配有无标签数据。

我们提出 VisualDeltas,一种轻量级偏好学习框架,通过多模态数据中图像质量的变化提取监督信号。利用图像质量对视觉感知和推理的系统性影响,VisualDeltas 在不依赖人类标注或外部教师的情况下,诱导出有信息量的偏好信号。该框架支持无标签和带标签两种模式,可在有监督信息可用时灵活使用。在多种多模态基准和模型规模下,VisualDeltas 均显著优于拒绝采样微调,提升模型泛化性能,并可自然扩展至多种视觉退化场景。

原文摘要 · Abstract (English)

We present VisualDeltas, a lightweight preference-learning framework that extracts supervision from visual quality variations in multimodal data. By leveraging the systematic impact of image quality on visual perception and reasoning, VisualDeltas induces informative preference signals without relying on human annotations or external teachers. The framework supports both label-free and label-based regimes, enabling flexible use of available supervision when present. Across diverse multimodal benchmarks and model scales, VisualDeltas consistently outperforms rejection-sampling fine-tuning and improves generalization, and extends naturally to a range of visual degradations.

偏好学习多模态无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。