arXiv:2411.17776cs.CVcs.MM2024-11ICCV被引 15

构建大规模图文数据集,实现基于文本的异常行为行人搜索。

Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search

  • 提出图文结合的异常行为行人检索任务,涵盖正常与异常动作。
  • 构建包含101万合成图文对的PAB数据集,测试集含1978个真实样本。
  • 引入姿态感知框架,召回率达84.93%,适合安防与监控场景应用。

文本驱动的行人搜索旨在通过自然语言描述跨摄像头网络检索特定个体。然而,现有基准常偏向于行走、站立等常见动作,忽视了现实场景中识别异常行为的关键需求。为此,我们提出新任务——基于文本的行人异常行为搜索,即通过文本定位执行常规或异常活动的行人。为支持该任务的训练与评估,我们构建了大规模图像-文本行人异常行为(PAB)基准,涵盖跑步、表演、踢足球等动作及其对应异常状态(如倒地、被击打、跌倒),同一身份下均有正常与异常实例。PAB训练集包含1,013,605个合成图像-文本对,测试集包含1,978个真实世界图像-文本对。为验证其有效性,我们提出一种跨模态姿态感知框架,融合人体姿态特征并采用基于身份的困难负样本采样。在所提基准上的大量实验表明,合成训练数据有助于细粒度行为检索,所提方法在recall@1上达到84.93%,优于其他竞争方法。数据集、模型及代码已开源。

原文摘要 · Abstract (English)

Text-based person search aims to retrieve specific individuals across camera networks using natural language descriptions. However, current benchmarks often exhibit biases towards common actions like walking or standing, neglecting the critical need for identifying abnormal behaviors in real-world scenarios. To meet such demands, we propose a new task, text-based person anomaly search, locating pedestrians engaged in both routine or anomalous activities via text. To enable the training and evaluation of this new task, we construct a large-scale image-text Pedestrian Anomaly Behavior (PAB) benchmark, featuring a broad spectrum of actions, e.g., running, performing, playing soccer, and the corresponding anomalies, e.g., lying, being hit, and falling of the same identity. The training set of PAB comprises 1,013,605 synthesized image-text pairs of both normalities and anomalies, while the test set includes 1,978 real-world image-text pairs. To validate the potential of PAB, we introduce a cross-modal pose-aware framework, which integrates human pose patterns with identity-based hard negative pair sampling. Extensive experiments on the proposed benchmark show that synthetic training data facilitates the fine-grained behavior retrieval, and the proposed pose-aware method arrives at 84.93% recall@1 accuracy, surpassing other competitive methods. The dataset, model, and code are available at https://github.com/Shuyu-XJTU/CMP.

行人搜索异常检测图文检索姿态感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。