arXiv:2510.02728cs.RO2025-10被引 6

用文字描述精准找图,让无人机在复杂空中场景中导航更准。

Team Xiaomi EV-AD VLA: Caption-Guided Retrieval System for Cross-Modal Drone Navigation -- Technical Report for IROS 2025 RoboSense Challenge Track 4

  • 先粗筛20张相关图,再用视觉语言模型生成详细描述,精细重排结果。
  • 在所有关键指标上提升5%,召回率显著优于基线方法。
  • 适合需要高精度跨模态检索的无人机、机器人导航场景。

跨模态无人机导航在机器人领域仍具挑战性,需基于自然语言描述从大规模数据库中高效检索相关图像。RoboSense 2025 Track 4挑战聚焦于多平台(无人机、卫星、地面摄像头)间鲁棒的自然语言引导跨视角图像检索。现有基线方法虽在初步检索中有效,但在复杂空域场景下难以实现文本与视觉内容间的细粒度语义匹配。为此,我们提出两阶段检索优化方法——标题引导检索系统(CGRS),通过智能重排序增强基线粗排序。首先利用基线模型获取每条查询的前20张最相关图像;随后使用视觉语言模型(VLM)为候选图像生成详细描述,捕捉其丰富语义信息;再基于多模态相似性计算框架,对原始文本查询进行细粒度重排序,构建视觉内容与自然语言之间的语义桥梁。本方法在各项关键指标(Recall@1、Recall@5、Recall@10)上均取得一致5%的提升,最终在挑战中获得第二名,验证了该语义精修策略在真实机器人导航场景中的实用价值。

原文摘要 · Abstract (English)

Cross-modal drone navigation remains a challenging task in robotics, requiring efficient retrieval of relevant images from large-scale databases based on natural language descriptions. The RoboSense 2025 Track 4 challenge addresses this challenge, focusing on robust, natural language-guided cross-view image retrieval across multiple platforms (drones, satellites, and ground cameras). Current baseline methods, while effective for initial retrieval, often struggle to achieve fine-grained semantic matching between text queries and visual content, especially in complex aerial scenes. To address this challenge, we propose a two-stage retrieval refinement method: Caption-Guided Retrieval System (CGRS) that enhances the baseline coarse ranking through intelligent reranking. Our method first leverages a baseline model to obtain an initial coarse ranking of the top 20 most relevant images for each query. We then use Vision-Language-Model (VLM) to generate detailed captions for these candidate images, capturing rich semantic descriptions of their visual content. These generated captions are then used in a multimodal similarity computation framework to perform fine-grained reranking of the original text query, effectively building a semantic bridge between the visual content and natural language descriptions. Our approach significantly improves upon the baseline, achieving a consistent 5\% improvement across all key metrics (Recall@1, Recall@5, and Recall@10). Our approach win TOP-2 in the challenge, demonstrating the practical value of our semantic refinement strategy in real-world robotic navigation scenarios.

跨模态检索无人机导航视觉语言模型重排序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。