用语言指令动态引导视觉编码器,实现可控且通用的感知。
Language-Instructed Vision Embeddings for Controllable and Generalizable Perception

- 用语言指令在推理时动态调整视觉特征提取
- 在视觉问答上超越参数多出数倍的模型,幻觉减少34分
- 支持未见过的任务指令,适合需要灵活感知的场景
视觉基础模型通常作为静态特征提取器,任务适配依赖大型下游模型。本文提出新范式:不只将视觉特征输入语言模型,而是用语言本身动态引导视觉编码器。所提方法Language-Instructed Vision Embeddings(LIVE)利用语言作为高层指导,在推理时生成以任务为中心的嵌入,无需针对具体任务重新训练。这使编码器聚焦于输入中与上下文相关的内容,产生更可控、更具泛化性的表示。实验表明,LIVE在MMVP上减少视觉幻觉达34分,视觉问答性能超越参数量大几个数量级的视觉语言模型,并能泛化至未见指令和任务,为实现可适应、指令驱动的视觉智能提供了直接路径。
原文摘要 · Abstract (English)
Vision foundation models are typically trained as static feature extractors, placing the burden of task adaptation onto large downstream models. We propose an alternative paradigm: instead of solely feeding visual features into language models, we use language itself to dynamically guide the vision encoder. Our method, Language-Instructed Vision Embeddings (LIVE), leverages language as high-level guidance to produce task-centric embeddings at inference time, removing the need for task-specific retraining. This enables the encoder to focus on contextually relevant aspects of the input, yielding more controllable and generalizable representations. Empirically, LIVE reduces visual hallucinations (+34 points on MMVP), surpasses vision-language models with orders of magnitude more parameters on visual question answering, and generalizes to unseen instructions and tasks -- offering a direct path toward adaptive, instruction-driven visual intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。