用语言提示增强视觉模型对遮挡和小物体的定位能力
Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues

- 从多模态大模型视觉流提取语义线索,结合文本嵌入生成语言引导语义先验
- 在拥挤场景中显著提升定位准确率,有效缓解遮挡与小物体带来的性能下降
- 适合需要鲁棒目标定位的自动驾驶、机器人视觉等实际应用
尽管多模态大语言模型(MLLMs)在一般场景中已提升定位能力,但在密集场景下的鲁棒性仍待探索。密集场景中的视觉挑战(如遮挡和小物体)会损害物体语义并降低定位性能。相比之下,语言表达不受此类退化影响,能保持物体语义。基于此,我们提出一种新方法,通过语言引导的语义线索(LGSCs)克服上述限制。具体而言,该方法引入语义线索提取器(SCE),从MLLM的视觉流水线中提取物体语义线索,并利用对应文本嵌入引导生成LGSCs作为语言语义先验。随后将这些线索重新整合至原始视觉流水线,以优化物体语义。大量实验与分析表明,将LGSCs融入MLLM可有效提升密集场景中的定位准确率。
原文摘要 · Abstract (English)
While Multimodal Large Language Models (MLLMs) have enhanced grounding capabilities in general scenes, their robustness in crowded scenes remains underexplored. Crowded scenes entail visual challenges (i.e., occlusion and small objects), which impair object semantics and degrade grounding performance. In contrast, language expressions are immune to such degradation and preserve object semantics. In light of these observations, we propose a novel method that overcomes such constraints by leveraging Language-Guided Semantic Cues (LGSCs). Specifically, our approach introduces a Semantic Cue Extractor (SCE) to derive semantic cues of objects from the visual pipeline of an MLLM. We then guide these cues using corresponding text embeddings to produce LGSCs as linguistic semantic priors. Subsequently, they are reintegrated into the original visual pipeline to refine object semantics. Extensive experiments and analyses demonstrate that incorporating LGSCs into an MLLM effectively improves grounding accuracy in crowded scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。