arXiv:2509.25528cs.CVcs.AI2025-09被引 1

用大模型提升户外场景中语言指代定位的准确性

LLM-RG: Referential Grounding in Outdoor Scenarios using Large Language Models

  • 结合视觉语言模型提取属性,大模型进行符号推理
  • 在Talk2Car上显著优于纯视觉或语言模型基线
  • 零样本应用,适合自动驾驶等复杂场景

户外驾驶场景中的指代定位因场景变化大、视觉相似物体多、动态元素复杂而困难。本文提出LLM-RG,一种混合式管道:利用现成的视觉语言模型提取细粒度属性,再通过大语言模型进行符号推理。该方法处理图像与自由形式指代表达时,先由大模型提取相关物体类型和属性,检测候选区域;再用视觉语言模型生成丰富视觉描述,并结合空间元数据形成自然语言提示,输入大模型进行链式思维推理,最终确定目标边界框。在Talk2Car基准测试中,LLM-RG显著优于基于大模型和视觉语言模型的基线。消融实验表明,加入3D空间线索可进一步提升定位效果。结果证明,零样本下视觉语言模型与大语言模型的互补优势,能实现鲁棒的户外指代定位。

原文摘要 · Abstract (English)

Referential grounding in outdoor driving scenes is challenging due to large scene variability, many visually similar objects, and dynamic elements that complicate resolving natural-language references (e.g., "the black car on the right"). We propose LLM-RG, a hybrid pipeline that combines off-the-shelf vision-language models for fine-grained attribute extraction with large language models for symbolic reasoning. LLM-RG processes an image and a free-form referring expression by using an LLM to extract relevant object types and attributes, detecting candidate regions, generating rich visual descriptors with a VLM, and then combining these descriptors with spatial metadata into natural-language prompts that are input to an LLM for chain-of-thought reasoning to identify the referent's bounding box. Evaluated on the Talk2Car benchmark, LLM-RG yields substantial gains over both LLM and VLM-based baselines. Additionally, our ablations show that adding 3D spatial cues further improves grounding. Our results demonstrate the complementary strengths of VLMs and LLMs, applied in a zero-shot manner, for robust outdoor referential grounding.

指代定位大模型自动驾驶视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。