探究文本如何影响图像质量判断,发现现有模型依赖图像而非文本。
Understanding Pure Textual Reasoning for Blind Image Quality Assessment
- 用三种新范式分析文本与图像质量评分的关系
- 仅用文本预测时模型性能大幅下降,证明文本信息不足
- 自洽性方法显著缩小图文预测差距,适合提升文本推理
文本推理近年被广泛用于盲图像质量评估(BIQA),但其如何贡献评分、文本能多大程度代表与评分相关的内容仍不明确。本文从信息流视角出发,对比现有模型与三种设计用于学习图像-文本-评分关系的新范式:链式思维(Chain-of-Thought)、自洽性(Self-Consistency)和自编码器(Autoencoder)。实验表明,当仅使用文本信息进行预测时,现有模型性能显著下降。链式思维范式对BIQA性能提升有限,而自洽性范式显著缩小了图像与文本条件下的预测差距,使PLCC/SRCC差异降至0.02/0.03。自编码器范式在弥合图文差距方面效果较弱,但揭示了优化方向。研究为改进文本推理在BIQA及高层视觉任务中的应用提供了洞见。
原文摘要 · Abstract (English)
Textual reasoning has recently been widely adopted in Blind Image Quality Assessment (BIQA). However, it remains unclear how textual information contributes to quality prediction and to what extent text can represent the score-related image contents. This work addresses these questions from an information-flow perspective by comparing existing BIQA models with three paradigms designed to learn the image-text-score relationship: Chain-of-Thought, Self-Consistency, and Autoencoder. Our experiments show that the score prediction performance of the existing model significantly drops when only textual information is used for prediction. Whereas the Chain-of-Thought paradigm introduces little improvement in BIQA performance, the Self-Consistency paradigm significantly reduces the gap between image- and text-conditioned predictions, narrowing the PLCC/SRCC difference to 0.02/0.03. The Autoencoder-like paradigm is less effective in closing the image-text gap, yet it reveals a direction for further optimization. These findings provide insights into how to improve the textual reasoning for BIQA and high-level vision tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。