提出首个空地跨视图指代检测基准与一致性定位框架
A-PAIR: A Benchmark and Identity-Consistent Grounding Framework for Air-Ground Cross-View Referring Person Detection

- 设计分因子标注与对齐机制,低成本构建跨视图指代表达数据集
- 引入联合空地配对检测与身份一致性校准,对齐准确率提升至22.28%
- 适用于多智能体协同任务中的语言-视觉-控制链路建模
空地跨视图指代人物检测是集体具身智能中语言到感知再到控制链条的关键环节,需在空中与地面代理间将语言指令精准定位到同一物理目标。现有指代理解与开放词汇定位方法未同时考虑跨视图身份一致性,难以应对相似行人干扰、弱空中外观线索及身份一致性要求。为此,本文提出首个综合性基准A-PAIR,包含22,137个跨视图指代表达样本。为高效构建该数据集,提出半自动标注框架FARA,实现因子化指代描述与身份一致性监督的生成。进一步提出身份一致指代定位(ICRG)框架,融合因子化指代定位、候选完整性监督与跨视图一致性校准,实现空地配对联合选择。ICRG在地面、空中及配对层面检测性能均优于强基线,配对F1从16.65%提升至22.28%。结果表明,空地跨视图指代检测需依赖配对检测与身份一致性推理。
原文摘要 · Abstract (English)
Air-ground cross-view referring person detection is a necessary component in the language-to-perception-to-control chain of collective embodied intelligence, grounding a language command into the same physical target before ground and aerial agents can coordinate downstream actions. Existing referring expression comprehension and open-vocabulary grounding methods do not jointly account for cross-view identity consistency, making them insufficient for Air-Ground Cross-View Referring Person Detection (AGCV-RPD), which involves similar pedestrian distractors, weak aerial appearance cues, and cross-view identity consistency. To study this problem, we introduce Air-Ground Paired Identity-Aware Referring (A-PAIR), the first comprehensive AGCV-RPD benchmark, containing 22,137 cross-view referring samples. To construct A-PAIR efficiently, we propose Factorized Annotation and Referential Alignment (FARA), a semi-automatic annotation framework that generates factorized referring descriptions and identity-consistency supervision at reduced cost. We propose Identity-Consistent Referring Grounding (ICRG), a framework that combines factorized referential grounding, candidate-completeness supervision, and cross-view consistency calibration for joint air-ground pair selection. ICRG improves ground, aerial, and pair-level detection over strong baselines, increasing pair F1 from 16.65% to 22.28%. These results show that AGCV-RPD requires paired detection and identity-consistent reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。