用贝叶斯方法让大模型更准地找高分辨率图里的小物体
BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception

- 把视觉搜索建模为连续空间尺度的全局优化问题
- 在超高清图像中检测精度显著优于现有方法
- 适合需要精细感知的医学影像、遥感等场景
虽然多模态大语言模型(MLLMs)展现出强大的通用能力,但在超高清(UHR)图像中对细粒度感知仍存在困难,尤其在杂乱场景中检测微小目标。现有方法面临两难:要么依赖低效的无先验扫描,要么依赖静态先验启发式策略,缺乏后验修正来纠正初始模型偏差。为此,我们提出BVS(贝叶斯视觉搜索),将感知建模为在连续空间-尺度流形上的全局优化问题。BVS融合先验引导与后验修正:利用MLLM的早停注意力传播构建推理感知先验,同时采用尺度感知的非平稳核与高斯过程-上置信界(GP-UCB)机制,通过迭代局部观测动态修正噪声并恢复缺失信息。理论上,我们给出了次线性后悔界保证;大量实验表明,BVS在准确率与效率之间实现更优平衡,显著超越当前最优基线。
原文摘要 · Abstract (English)
While Multimodal Large Language Models (MLLMs) demonstrate impressive general capabilities, they struggle with fine-grained perception in ultra-high-resolution (UHR) images, particularly for tiny objects in cluttered scenes. Existing methods face a dilemma: they either rely on inefficient prior-free scanning, or depend on static prior-driven heuristics that lack posterior correction to rectify initial model biases. To address this, we propose BVS (Bayesian Visual Search), a framework that formulates perception as a global optimization problem over a continuous spatial-scale manifold. Specifically, BVS bridges prior guidance with posterior correction: it utilizes an early-stop attention rollout of MLLM to construct reasoning-aware priors, while employing a scale-aware non-stationary kernel and GP-UCB to dynamically rectify noise and recover missing information in the prior through iterative local observations. We provide theoretical guarantees via sub-linear regret bounds, and extensive experiments demonstrate that BVS significantly outperforms state-of-the-art baselines with a superior trade-off between accuracy and efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。