用感官提示让纯文本模型具备视觉听觉感知能力
Words That Make Language Models Perceive
- 通过'看见'、'听见'等提示激活模型潜在多模态表征
- 轻量级提示即可使文本模型与视觉/音频编码器对齐
- 适合想低成本提升模型感知能力的研究者
纯文本训练的大语言模型看似缺乏直接感知经验,但其内部表征隐式受到语言中多模态规律的影响。我们检验假设:显式感官提示可揭示这一潜在结构,使纯文本 LLM 的表征更接近专业视觉和音频编码器。当提示模型‘看见’或‘听见’时,它会将下一个词的预测视为基于未实际提供的视觉或听觉证据的条件生成。结果表明,轻量级提示工程可稳定激活纯文本训练模型中的模态适配表征。
原文摘要 · Abstract (English)
Large language models (LLMs) trained purely on text ostensibly lack any direct perceptual experience, yet their internal representations are implicitly shaped by multimodal regularities encoded in language. We test the hypothesis that explicit sensory prompting can surface this latent structure, bringing a text-only LLM into closer representational alignment with specialist vision and audio encoders. When a sensory prompt tells the model to 'see' or 'hear', it cues the model to resolve its next-token predictions as if they were conditioned on latent visual or auditory evidence that is never actually supplied. Our findings reveal that lightweight prompt engineering can reliably activate modality-appropriate representations in purely text-trained LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。