arXiv:2608.23503cs.CV2026-08中稿 · ECCV

通过动作对齐检索与成对多模态重排,提升文本驱动行人异常行为搜索精度

Action-Aligned Retrieval with Pairwise Multimodal Reranking for Text-Based Person Anomaly Search

论文配图:Action-Aligned Retrieval with Pairwise Multimodal Reranking for Text-Based Person Anomaly Search
图 1 · 摘自论文原文
  • 三阶段框架:动作对齐微调+双路语义检索+高效成对重排
  • 在PAB数据集上超越现有方法,跨数据集迁移效果良好
  • 适合需要细粒度行为理解的视频监控场景

文本驱动行人异常行为搜索需基于细微且依赖上下文的行为差异进行个体区分,而非仅依赖外观。现有方法难以捕捉上下文相关的动作特征,常依赖孤立骨骼几何,忽略原始查询细节,或采用绝对点对评分进行多模态验证。为此,我们提出 extbf{ActPair},一种统一的粗到精三阶段框架,结合动作对齐检索与成对多模态重排,弥合姿态-语义鸿沟。首先,使用动作对齐多任务目标微调视觉语言模型(VLM),使表征编码动作区分性语义。其次,利用原始查询与大语言模型(LLM)生成的上下文增强重写版本并行晚期融合检索,保留双重语义互补信息。最后,提出一种高效现成重排模块,采用枢轴促进算法实现直接成对视觉对比,缓解残余空间与构图模糊问题,无需消耗巨大的全量评估开销。大量实验表明,该框架在公开的行人异常行为(PAB)测试集上表现最优,并能有效迁移至未见过的非异常特定数据集。

原文摘要 · Abstract (English)

Text-based person anomaly search requires distinguishing individuals based on fine-grained, context-dependent behaviors rather than mere appearance. Existing methods struggle to capture these context-conditioned actions, frequently relying on isolated skeletal geometry, discarding raw query details during reformulation, or utilizing absolute pointwise scoring for multimodal verification. To address these limitations, we propose \textbf{ActPair}, a unified three-stage coarse-to-fine framework that combines action-aligned retrieval with pairwise multimodal reranking to bridge the pose-semantic gap. First, we fine-tune a vision-language model (VLM) with an action-aligned multi-task objective that encourages the representations to encode action-discriminative semantics. Second, we perform parallel late-fusion retrieval using the original query and a large language model (LLM)-generated context-grounded rewrite, retaining complementary details from both semantic views. Finally, we propose an efficient off-the-shelf reranking module that leverages a pivot-promote algorithm to perform direct pairwise visual comparisons, mitigating residual spatial and compositional ambiguities without the prohibitive inference costs of exhaustive evaluation. Extensive experiments demonstrate that our framework achieves the best results among the compared methods on the Pedestrian Anomaly Behavior (PAB) public test and transfers effectively to an unseen, non-anomaly-specific dataset.

异常检测多模态检索文本搜索视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。