arXiv:2510.11369cs.CV2025-10被引 18

用对比学习让图像直接对齐文本表示,大幅降低视觉强化学习模型的开销。

Reasoning as Representation: Rethinking Visual Reinforcement Learning in Image Quality Assessment

  • 通过对比学习将图像直接对齐推理生成的通用文本表示
  • 性能接近推理型模型,参数量低于5%,推理时间也大幅减少
  • 适合资源受限场景下的图像质量评估应用

基于推理的图像质量评估(IQA)模型通过强化学习(RL)训练展现出卓越的泛化能力,但其内在机制与关键驱动因素尚未充分研究。尽管性能优异,这些模型的推理能耗和延迟比早期模型高出数个数量级,限制了实际部署。本文通过大量实验验证并阐明:通过强化学习训练,多模态大语言模型(MLLMs)利用其推理能力,将冗余的视觉表征转化为紧凑且跨域对齐的文本表征,这正是其泛化能力的来源。基于此发现,我们提出新算法RALI,采用对比学习直接将图像与RL学习到的通用文本表示对齐,无需依赖推理过程,甚至无需加载大语言模型(LLM)。在质量评分任务中,该框架性能接近基于推理的模型,同时模型参数少于5%,推理时间也显著降低。

原文摘要 · Abstract (English)

Reasoning-based image quality assessment (IQA) models trained through reinforcement learning (RL) exhibit exceptional generalization, yet the underlying mechanisms and critical factors driving this capability remain underexplored in current research. Moreover, despite their superior performance, these models incur inference energy usage and latency orders of magnitude higher than their earlier counterparts, restricting their deployment in specific scenarios. Through extensive experiments, this paper verifies and elaborates that through RL training, MLLMs leverage their reasoning capability to convert redundant visual representations into compact, cross-domain aligned text representations. This conversion is precisely the source of the generalization exhibited by these reasoning-based IQA models. Building on this fundamental insight, we propose a novel algorithm, RALI, which employs contrastive learning to directly align images with these generalizable text representations learned by RL. This approach eliminates the reliance on reasoning processes and even obviates the need to load an LLM. For the quality scoring task, this framework achieves generalization performance comparable to reasoning-based models while requiring less than 5% of their model parameters and inference time.

视觉强化学习图像质量评估对比学习轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。