arXiv:2511.08369cs.CVcs.AI2025-11AAAI被引 3

用文字描述跨视角检索空中与地面人像,解决视角差异难题。

Text-based Aerial-Ground Person Retrieval

  • 设计分层专家路由模块,分离视角特异性与通用特征。
  • 在新数据集上达到78.3%的mAP,优于现有方法。
  • 适合智能监控、搜救等跨视角图像检索场景。

本文提出文本驱动的跨视角空中-地面行人重识别(TAG-PR),旨在通过文本描述从异构的空中与地面视图中检索行人图像。与仅关注地面视角的传统文本行人重识别不同,该任务因视角差异大而更具实际意义且挑战显著。为此,我们构建了TAG-PEDES数据集,基于公开基准自动生成文本描述,并采用多样化的文本生成范式以增强在视角异质性下的鲁棒性;同时提出TAG-CLIP框架,通过分层路由专家模块学习视角特定与视角无关特征,并引入视角解耦策略以提升跨模态对齐效果。我们在新提出的TAG-PEDES数据集及现有T-PR基准上评估了该方法的有效性。代码与数据集已开源。

原文摘要 · Abstract (English)

This work introduces Text-based Aerial-Ground Person Retrieval (TAG-PR), which aims to retrieve person images from heterogeneous aerial and ground views with textual descriptions. Unlike traditional Text-based Person Retrieval (T-PR), which focuses solely on ground-view images, TAG-PR introduces greater practical significance and presents unique challenges due to the large viewpoint discrepancy across images. To support this task, we contribute: (1) TAG-PEDES dataset, constructed from public benchmarks with automatically generated textual descriptions, enhanced by a diversified text generation paradigm to ensure robustness under view heterogeneity; and (2) TAG-CLIP, a novel retrieval framework that addresses view heterogeneity through a hierarchically-routed mixture of experts module to learn view-specific and view-agnostic features and a viewpoint decoupling strategy to decouple view-specific features for better cross-modal alignment. We evaluate the effectiveness of TAG-CLIP on both the proposed TAG-PEDES dataset and existing T-PR benchmarks. The dataset and code are available at https://github.com/Flame-Chasers/TAG-PR.

跨视角检索文本检索多模态行人重识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。