arXiv:2603.20721cs.CV2026-03中稿 · CVPR

用模糊逻辑提升文本与航拍图像的细粒度对齐,增强识别鲁棒性。

Cross-modal Fuzzy Alignment Network for Text-Aerial Person Retrieval and A Large-scale Benchmark

  • 引入模糊逻辑动态评估词元可靠性,抑制噪声干扰。
  • 通过地面视图作为桥梁,减少航拍图像与文本的语义差距。
  • 构建大规模数据集AERI-PEDES,提升文本描述准确性和一致性。

文本-航拍人物检索旨在从无人机拍摄的图像中根据目击者描述定位目标,支持智能交通与公共安全应用。相比地面视角的文本-图像检索,无人机图像常因视角和飞行高度剧烈变化导致视觉信息退化,使文本与图像语义对齐极具挑战。为此,我们提出一种新型跨模态模糊对齐网络,通过模糊逻辑量化词元级可靠性,实现精准细粒度对齐,并引入地面视图作为桥梁代理,进一步缩小航拍图像与文本描述之间的差距。具体地,设计了模糊词元对齐模块,利用模糊隶属函数动态建模词元关联强度,抑制不可见或噪声词元的影响,缓解因视觉线索缺失导致的语义不一致,显著提升词元级对齐鲁棒性。此外,设计上下文感知动态对齐模块,将地面视图作为桥梁,在直接对齐与代理辅助对齐间自适应融合,进一步提升性能。同时,构建大规模基准数据集AERI-PEDES,采用链式思维分解文本生成为属性解析、初始描述与精炼三步,显著提升文本准确性与语义一致性。在AERI-PEDES与TBAPR上的实验验证了方法优越性。

原文摘要 · Abstract (English)

Text-aerial person retrieval aims to identify targets in UAV-captured images from eyewitness descriptions, supporting intelligent transportation and public security applications. Compared to ground-view text--image person retrieval, UAV-captured images often suffer from degraded visual information due to drastic variations in viewing angles and flight altitudes, making semantic alignment with textual descriptions very challenging. To address this issue, we propose a novel Cross-modal Fuzzy Alignment Network, which quantifies the token-level reliability by fuzzy logic to achieve accurate fine-grained alignment and incorporates ground-view images as a bridge agent to further mitigate the gap between aerial images and text descriptions, for text--aerial person retrieval. In particular, we design the Fuzzy Token Alignment module that employs the fuzzy membership function to dynamically model token-level association strength and suppress the influence of unobservable or noisy tokens. It can alleviate the semantic inconsistencies caused by missing visual cues and significantly enhance the robustness of token-level semantic alignment. Moreover, to further mitigate the gap between aerial images and text descriptions, we design a Context-Aware Dynamic Alignment module to incorporate the ground-view agent as a bridge in text--aerial alignment and adaptively combine direct alignment and agent-assisted alignment to improve the robustness. In addition, we construct a large-scale benchmark dataset called AERI-PEDES by using a chain-of-thought to decompose text generation into attribute parsing, initial captioning, and refinement, thus boosting textual accuracy and semantic consistency. Experiments on AERI-PEDES and TBAPR demonstrate the superiority of our method.

文本检索航拍图像模糊逻辑多模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。