让视觉语言模型像侦探一样主动选择要看的细节,提升复杂推理能力。
The Perceptual Bandwidth Bottleneck in Vision-Language Models: Active Visual Reasoning via Sequential Experimental Design

- 用序列实验设计主动挑选关键视觉信息,突破感知带宽限制。
- 在高分辨率遥感任务中,性能超越直接推理和ReAct方法。
- 无需训练,通过证据导向探测优化图像裁剪,适合精细视觉任务。
现代视觉语言模型(VLMs)的视觉感知受限于感知带宽瓶颈:宽视野虽保留全局上下文,却牺牲了复杂推理所需的细粒度细节。我们认为,高分辨率视觉推理不仅是语义理解,更是有限感知带宽下对任务相关证据的主动获取。受主动视觉和信息觅食启发,我们将该过程形式化为序列贝叶斯最优实验设计(S-BOED),即智能体在回答前决定采集哪些视觉证据。由于连续吉像素空间中的精确贝叶斯推断不可行,我们推导出一个可计算的覆盖-分辨率目标,作为任务相关信息增益的代理。我们基于此构建了FOVEA,一种无需训练的流程,通过证据导向探测优化VLM的图像裁剪提案。在高分辨率基准测试中,相比直接推理和ReAct风格基线,性能持续提升,尤其在以搜索为主的遥感场景中表现突出。
原文摘要 · Abstract (English)
Visual perception in modern Vision-Language Models (VLMs) is constrained by a perceptual bandwidth bottleneck: a broad field of view preserves global context but sacrifices the fine-grained details required for complex reasoning. We argue that high-resolution visual reasoning is therefore not only semantic reasoning but also task-relevant evidence acquisition under limited perceptual bandwidth. Inspired by active vision and information foraging, we formalise this process as sequential Bayesian optimal experimental design (S-BOED), where an agent decides which visual evidence to acquire before answering. Since exact Bayesian inference is intractable in continuous gigapixel spaces, we derive a tractable coverage--resolution objective as a proxy for task-relevant information gain. We instantiate this framework with FOVEA, a training-free procedure that refines VLM crop proposals through evidence-oriented probing. Experiments on high-resolution benchmarks show consistent gains over direct and ReAct-style baselines, with particularly strong improvements in search-dominated remote-sensing settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。