arXiv:2603.22641cs.CV2026-03

让AI在隐空间推理图像质量,比用语言更高效

Q-Tacit: Image Quality Assessment via Latent Visual Reasoning

  • 在隐空间注入视觉质量先验,让模型直接思考图像质量
  • 用更少的词(token)完成高质量评估,性能更强
  • 适合做视觉质量评估、想突破语言瓶颈的研究者

基于视觉-语言模型(VLM)的图像质量评估(IQA)通过引入思维链(CoT)推理取得了显著进展。近期工作通过强化学习和主动视觉工具优化了质量推理,但这些方法仍以语言为中心,将视觉信息视为静态前提。由于离散文本标记与质量感知空间之间存在鸿沟,许多质量相关的视觉线索难以完整转译为语言,限制了复杂视觉任务中的推理效果。本文提出新范式Q-Tacit,重新思考‘自然语言是否是质量推理的理想空间’这一问题,引导VLM在隐空间中进行超越语言的视觉质量推理。该方法采用双阶段协同流程:(i) 将结构化视觉质量先验注入隐空间,(ii) 校准隐空间推理轨迹以提升评估能力。大量实验表明,Q-Tacit在使用远少于以往方法的词元(tokens)情况下,仍实现优异的整体性能。本研究验证了语言并非唯一适用于视觉质量的紧凑表征,为未来探索高效的隐空间推理范式开辟了新路径。源代码将公开以支持后续研究。

原文摘要 · Abstract (English)

Vision-Language Model (VLM)-based image quality assessment (IQA) has been significantly advanced by incorporating Chain-of-Thought (CoT) reasoning. Recent work has refined image quality reasoning by applying reinforcement learning (RL) and leveraging active visual tools. However, such strategies are typically language-centric, with visual information being treated as static preconditions. Quality-related visual cues often cannot be abstracted into text in extenso due to the gap between discrete textual tokens and quality perception space, which in turn restricts the reasoning effectiveness for visually intensive IQA tasks. In this paper, we revisit this by asking the question, "Is natural language the ideal space for quality reasoning?" and, as a consequence, we propose Q-Tacit, a new paradigm that elicits VLMs to reason beyond natural language in the latent quality space. Our approach follows a synergistic two-stage process: (i) injecting structural visual quality priors into the latent space, and (ii) calibrating latent reasoning trajectories to improve quality assessment ability. Extensive experiments demonstrate that Q-Tacit can effectively perform quality reasoning with significantly fewer tokens than previous reasoning-based methods, while achieving strong overall performance. This paper validates the proposition that language is not the only compact representation suitable for visual quality, opening possibilities for further exploration of effective latent reasoning paradigms for IQA. Source code will be released to support future research.

图像质量评估隐空间推理视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。