针对无人机视频文本检索难题,提出多语义自适应挖掘方法
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
- 设计多语义自适应学习机制,动态捕捉帧间变化与关键区域语义
- 在两个自建数据集上超越现有方法,显著提升检索准确率
- 适合关注无人机视觉理解与跨模态检索的研究者
随着无人机技术发展,视频数据量急剧增加,亟需高效的语义检索方案。本文首次系统性地提出并研究无人机视频-文本检索(DVTR)任务。无人机视频具有俯视视角、强结构同质性以及目标组合的多样化语义表达等特点,挑战了为地面视角设计的现有跨模态方法对这些特性的建模能力。为此,本文提出多语义自适应挖掘(MSAM)方法,引入多语义自适应学习机制,结合帧间动态变化,从特定场景区域提取丰富语义信息,增强对无人机视频内容的深层理解与推理。该方法通过细粒度的文本-视频帧交互,融合自适应语义构建模块、分布驱动的语义学习项和多样性语义项,深化文本与视频模态间的交互,提升特征表示鲁棒性。为降低复杂背景干扰,引入跨模态交互特征融合池化机制,聚焦目标区域的特征提取与匹配,减少噪声影响。在两个自建无人机视频-文本数据集上的大量实验表明,MSAM在无人机视频-文本检索任务中优于现有方法。源代码与数据集将公开发布。
原文摘要 · Abstract (English)
With the advancement of drone technology, the volume of video data increases rapidly, creating an urgent need for efficient semantic retrieval. We are the first to systematically propose and study the drone video-text retrieval (DVTR) task. Drone videos feature overhead perspectives, strong structural homogeneity, and diverse semantic expressions of target combinations, which challenge existing cross-modal methods designed for ground-level views in effectively modeling their characteristics. Therefore, dedicated retrieval mechanisms tailored for drone scenarios are necessary. To address this issue, we propose a novel approach called Multi-Semantic Adaptive Mining (MSAM). MSAM introduces a multi-semantic adaptive learning mechanism, which incorporates dynamic changes between frames and extracts rich semantic information from specific scene regions, thereby enhancing the deep understanding and reasoning of drone video content. This method relies on fine-grained interactions between words and drone video frames, integrating an adaptive semantic construction module, a distribution-driven semantic learning term and a diversity semantic term to deepen the interaction between text and drone video modalities and improve the robustness of feature representation. To reduce the interference of complex backgrounds in drone videos, we introduce a cross-modal interactive feature fusion pooling mechanism that focuses on feature extraction and matching in target regions, minimizing noise effects. Extensive experiments on two self-constructed drone video-text datasets show that MSAM outperforms other existing methods in the drone video-text retrieval task. The source code and dataset will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。