针对无人机视频设计细粒度图文检索数据集与对齐方法
TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval
- 通过多粒度对齐实现文本与视频的精准匹配
- 在新数据集上达45.5%文本到视频检索准确率
- 适合无人机视觉理解与跨模态检索研究者
无人机已成为实时高分辨率数据采集的重要平台,产生海量航拍视频。高效检索相关视频内容对城市管理、应急响应、安防和救灾至关重要。尽管自然视频领域的图文检索已有进展,但无人机领域因现有数据集存在描述粗略、冗余等问题而研究不足。为此,本文构建了包含2,864段视频和14,320条细粒度、语义多样标注的无人机视频-文本匹配数据集(DVTMD),其标注涵盖人类行为、物体、背景环境、气象条件及视觉风格等多方面信息,增强图文对应性并降低冗余。基于此数据集,提出文本引导的多粒度对齐(TCMA)框架,融合全局视频-句子对齐、句子引导帧聚合与词引导块对齐。为优化局部对齐,设计词与块选择模块过滤无关内容,并引入文本自适应动态温度机制以根据文本类型调节注意力锐度。在DVTMD和CapERA上的大量实验建立了首个完整的无人机图文检索基准。所提TCMA方法在文本到视频检索中达到45.5% R@1,在视频到文本检索中达42.8% R@1,验证了数据集与方法的有效性。代码与数据集将公开。
原文摘要 · Abstract (English)
Unmanned aerial vehicles (UAVs) have become powerful platforms for real-time, high-resolution data collection, producing massive volumes of aerial videos. Efficient retrieval of relevant content from these videos is crucial for applications in urban management, emergency response, security, and disaster relief. While text-video retrieval has advanced in natural video domains, the UAV domain remains underexplored due to limitations in existing datasets, such as coarse and redundant captions. Thus, in this work, we construct the Drone Video-Text Match Dataset (DVTMD), which contains 2,864 videos and 14,320 fine-grained, semantically diverse captions. The annotations capture multiple complementary aspects, including human actions, objects, background settings, environmental conditions, and visual style, thereby enhancing text-video correspondence and reducing redundancy. Building on this dataset, we propose the Text-Conditioned Multi-granularity Alignment (TCMA) framework, which integrates global video-sentence alignment, sentence-guided frame aggregation, and word-guided patch alignment. To further refine local alignment, we design a Word and Patch Selection module that filters irrelevant content, as well as a Text-Adaptive Dynamic Temperature Mechanism that adapts attention sharpness to text type. Extensive experiments on DVTMD and CapERA establish the first complete benchmark for drone text-video retrieval. Our TCMA achieves state-of-the-art performance, including 45.5% R@1 in text-to-video and 42.8% R@1 in video-to-text retrieval, demonstrating the effectiveness of our dataset and method. The code and dataset will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。