通过多源信息融合提升罕见异常检测与分类准确率
Anomize: Better Open Vocabulary Video Anomaly Detection
- 融合视觉与文本信息,增强对未知异常的判别能力
- 利用标签间关系引导编码,减少新异常误分类
- 在UCF-Crime和XD-Violence上表现领先,适合开放词汇异常检测
开放词汇视频异常检测(OVVAD)旨在识别并分类基础异常与新型异常。然而现有方法在处理新型异常时面临两大挑战:一是检测模糊性,模型难以对不熟悉异常赋予准确异常得分;二是分类混淆,新型异常常被错误归类为视觉相似的基础类别。为此,本文从多源补充信息出发,结合多层次视觉数据与匹配文本信息,缓解检测模糊性;同时引入标签关系以指导新标签编码,增强新型视频与其对应标签的对齐,从而降低分类混淆。所提出的Anomize框架有效应对上述问题,在UCF-Crime与XD-Violence数据集上取得优异性能,验证了其在OVVAD任务中的有效性。
原文摘要 · Abstract (English)
Open Vocabulary Video Anomaly Detection (OVVAD) seeks to detect and classify both base and novel anomalies. However, existing methods face two specific challenges related to novel anomalies. The first challenge is detection ambiguity, where the model struggles to assign accurate anomaly scores to unfamiliar anomalies. The second challenge is categorization confusion, where novel anomalies are often misclassified as visually similar base instances. To address these challenges, we explore supplementary information from multiple sources to mitigate detection ambiguity by leveraging multiple levels of visual data alongside matching textual information. Furthermore, we propose incorporating label relations to guide the encoding of new labels, thereby improving alignment between novel videos and their corresponding labels, which helps reduce categorization confusion. The resulting Anomize framework effectively tackles these issues, achieving superior performance on UCF-Crime and XD-Violence datasets, demonstrating its effectiveness in OVVAD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。