arXiv:2604.23282cs.CVcs.MM2026-04ACL被引 5

用分阶段框架解决行为搜索中的动作语义歧义问题。

Bridging the Pose-Semantic Gap: A Cascade Framework for Text-Based Person Anomaly Search

论文配图:Bridging the Pose-Semantic Gap: A Cascade Framework for Text-Based Person Anomaly Search
图 1 · 摘自论文原文
  • 分两阶段检索:先按骨骼结构粗筛,再用多智能体验证语义。
  • 在PAB数据集上达到当前最优性能,兼顾效率与语义理解。
  • 适合需要精准文本检索的安防场景,尤其关注动作语义差异。

基于文本的人体异常搜索从监控视频中通过自然语言查询定位特定行为事件。尽管近期姿态感知方法能较好对齐几何结构,但仍存在根本性的姿态-语义鸿沟:语义不同的动作可能具有相似的骨架形态。虽然多模态大语言模型(MLLM)可缓解这种歧义,但用于大规模检索计算成本过高。本文提出结构-语义解耦级联框架(SSDC),将检索分为两步:(1) 基于骨架相似性的轻量级模型快速进行粗筛选;(2) 多智能体语义验证模块——由侦探(快速二分类)、分析师(提取证据)和写作者(生成语义摘要)组成。最终融合合成描述与结构先验进行重排序。在PAB基准测试中,SSDC实现了领先性能,平衡了效率与语义推理能力。

原文摘要 · Abstract (English)

Text-based person anomaly search retrieves specific behavioral events from surveillance archives using natural-language queries. Although recent pose-aware methods align geometric structures well, they face a fundamental Pose-Semantic Gap: semantically different actions can share similar skeletal geometries. While Multimodal Large Language Models (MLLMs) can reduce this ambiguity, using them for large-scale retrieval is computationally prohibitive. We propose the Structure-Semantic Decoupled Cascade (SSDC) framework, which decouples retrieval into two stages: (1) Structure-Aware Coarse Retrieval, where a lightweight model quickly filters candidates by skeletal similarity ; and (2) Detective Squad Interaction, a multi-agent semantic verification module. The squad consists of a Detective for fast binary filtering, an Analyst for evidence extraction, and a Writer for semantic synthesis. Finally, we re-rank candidates by fusing the synthesized captions with structural priors. Experiments on the PAB benchmark show that SSDC achieves state-of-the-art performance by balancing efficiency and semantic reasoning.

行为搜索多智能体语义鸿沟监控分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。