通过融合多模态模型并智能重排序,提升文本描述的异常行人检索准确率。
Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval

- 融合异构视觉语言模型,通过分数对齐逐步集成
- 在PAB数据集上达到90.92% mAP和98.68% Recall@10
- 适合需要精准异常行为识别的智慧安防场景
文本驱动的行人异常检索旨在利用自然语言描述从大规模图像库中找出具有异常行为的行人。与传统文本行人检索相比,该任务需对行人外观、行为、物体交互及场景上下文进行细粒度推理,跨模态匹配更具挑战性。本文提出GENAI4E团队参与AI City Challenge 2026 Track 4的解决方案:在强基线检索框架基础上,通过分数对齐与迭代集成融合异构视觉-语言嵌入模型,并引入基于分歧感知的VLM重排序机制处理模糊查询。在官方行人异常行为(PAB)基准上,本方法取得90.92% mAP、85.13% Recall@1、97.72% Recall@5和98.68% Recall@10的性能,验证了结合互补视觉-语言表征与选择性多模态推理在大规模文本驱动行人异常检索中的有效性。
原文摘要 · Abstract (English)
Text-based person anomaly retrieval aims to retrieve pedestrians exhibiting anomalous behaviors from a large image gallery using natural language descriptions. Compared with conventional text-based person retrieval, this task requires fine-grained reasoning over pedestrian appearance, behaviors, object interactions, and scene context, making robust cross-modal matching significantly more challenging. This paper presents the GENAI4E team's solution to AI City Challenge 2026 Track 4. Our framework builds upon a strong retrieval backbone and progressively integrates heterogeneous vision-language embedding models through score alignment and iterative ensemble fusion, followed by disagreement-aware VLM reranking for ambiguous queries. On the official Pedestrian Anomaly Behavior (PAB) benchmark, our approach achieves 90.92% mAP, 85.13% Recall@1, 97.72% Recall@5, and 98.68% Recall@10, demonstrating the effectiveness of combining complementary vision-language representations with selective multimodal reasoning for large-scale text-based person anomaly retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。