用生成图像激活文本模型的视觉先验,提升理解能力。
Synthetic Perception: Can Generated Images Unlock Latent Visual Prior for Text-Centric Reasoning?
- 用T2I模型实时生成图像,把文本映射到视觉语义空间。
- 在多个任务上显著提升Llama-3、Qwen等大模型性能。
- 适合缺乏视觉数据的文本任务,需关注图文语义对齐。
文本数据丰富但多模态模型日益强大之间存在显著的“模态鸿沟”。本文系统研究了由文生图(T2I)模型即时生成的图像是否能作为解锁文本中心推理中潜在视觉先验的机制。通过在文本分类任务上构建全面评估框架,分析了T2I模型质量(如Flux.1、SDXL)、提示工程策略及多模态融合架构等关键变量的影响。结果表明,这种“合成感知”可显著提升性能,有效将文本投影至视觉语义空间,即使在增强如Llama-3和Qwen-2.5等强语言模型基线时亦然。该方法可视为一种跨模态探测,缓解纯文本训练中的感官缺失问题。但效果高度依赖于文本与生成图像的语义对齐度、任务的视觉可接地性以及T2I模型生成保真度。本工作建立了该范式的严格基准,证明其在传统单模态场景中丰富语言理解的可行性。
原文摘要 · Abstract (English)
A significant ``modality gap" exists between the abundance of text-only data and the increasing power of multimodal models. This work systematically investigates whether images generated on-the-fly by Text-to-Image (T2I) models can serve as a mechanism to unlock latent visual priors for text-centric reasoning. Through a comprehensive evaluation framework on text classification, we analyze the impact of critical variables, including T2I model quality (e.g., Flux.1, SDXL), prompt engineering strategies, and multimodal fusion architectures. Our findings demonstrate that this ``synthetic perception" can yield significant performance gains by effectively projecting text into a visual semantic space, even when augmenting strong large language model baselines like Llama-3 and Qwen-2.5. We show that this approach serves as a form of cross-modal probing, mitigating the sensory deprivation inherent in pure text training. However, the effectiveness is highly conditional, depending on the semantic alignment between text and the generated image, the task's visual groundability, and the generative fidelity of the T2I model. Our work establishes a rigorous benchmark for this paradigm, demonstrating its viability as a pathway to enrich language understanding in traditionally unimodal scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。