arXiv:2502.03230cs.CVcs.MM2025-02中稿 · 2025 WWW Workshop …被引 4

针对文本描述相似的异常行人搜索,提出新匹配策略提升准确率。

Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search

  • 引入相似性覆盖分析策略,增强对文本细微差别的识别能力。
  • 在大规模图像库中实现高精度异常行人检索,显著提升搜索可靠性。
  • 适合需要精准图文匹配的智能安防与视频分析场景。

本文介绍HFUT-LMC团队在WWW 2025文本驱动行人异常搜索(TPAS)挑战赛中的解决方案。该任务旨在从大规模行人图像库中精准识别出具有正常或异常行为的行人。与传统视频分析不同,TPAS高度依赖模型对文本描述与视觉数据之间微妙关系的理解。其难点在于模型需在海量图像中将个体与文本精准匹配,并在描述相近时准确区分搜索结果。为此,我们提出相似性覆盖分析(SCA)策略,有效缓解因文本描述相似带来的识别困难,显著提升了模型在细微差异下的判别能力,从而增强搜索的准确性和可靠性。本方案在挑战赛中表现优异。

原文摘要 · Abstract (English)

This paper presents the HFUT-LMC team's solution to the WWW 2025 challenge on Text-based Person Anomaly Search (TPAS). The primary objective of this challenge is to accurately identify pedestrians exhibiting either normal or abnormal behavior within a large library of pedestrian images. Unlike traditional video analysis tasks, TPAS significantly emphasizes understanding and interpreting the subtle relationships between text descriptions and visual data. The complexity of this task lies in the model's need to not only match individuals to text descriptions in massive image datasets but also accurately differentiate between search results when faced with similar descriptions. To overcome these challenges, we introduce the Similarity Coverage Analysis (SCA) strategy to address the recognition difficulty caused by similar text descriptions. This strategy effectively enhances the model's capacity to manage subtle differences, thus improving both the accuracy and reliability of the search. Our proposed solution demonstrated excellent performance in this challenge.

异常检测图文匹配行人搜索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。