用专家模型融合解决多平台图像定位难题,精准匹配语言查询与异构影像。
A Parameter-Efficient Mixture-of-Experts Framework for Cross-Modal Geo-Localization
- 分平台处理+文本优化,对齐不同视角的图文语义差异。
- 三专家联合训练,硬负样本挖掘提升定位准确率,领先官方榜单。
- 适合跨模态检索、无人机导航等实际场景应用。
我们提出RoboSense 2025 Track 4:跨模态无人机导航任务的优胜方案。任务目标是根据自然语言查询,在大规模多平台图像库(卫星/无人机/地面)中检索最相关的地理参考图像。主要挑战来自平台间显著的异质性以及通用训练描述与平台特定测试查询之间的领域差距。我们通过领域对齐预处理流程与混合专家(MoE)框架缓解该问题:(i) 平台划分、卫星图像增强、去除方向词;(ii) 基于LLM的标题精炼管道,使文本语义与各平台视觉特征对齐。采用BGE-M3(文本)与EVA-CLIP(图像),以渐进式两阶段、硬负样本挖掘策略训练三个平台专家,并在推理时融合其得分。系统登顶官方排行榜,证明在异构视角下具备鲁棒的跨模态地理定位能力。
原文摘要 · Abstract (English)
We present a winning solution to RoboSense 2025 Track 4: Cross-Modal Drone Navigation. The task retrieves the most relevant geo-referenced image from a large multi-platform corpus (satellite/drone/ground) given a natural-language query. Two obstacles are severe inter-platform heterogeneity and a domain gap between generic training descriptions and platform-specific test queries. We mitigate these with a domain-aligned preprocessing pipeline and a Mixture-of-Experts (MoE) framework: (i) platform-wise partitioning, satellite augmentation, and removal of orientation words; (ii) an LLM-based caption refinement pipeline to align textual semantics with the distinct visual characteristics of each platform. Using BGE-M3 (text) and EVA-CLIP (image), we train three platform experts using a progressive two-stage, hard-negative mining strategy to enhance discriminative power, and fuse their scores at inference. The system tops the official leaderboard, demonstrating robust cross-modal geo-localization under heterogeneous viewpoints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。