用异常思维链让大模型更懂犯罪检测,效果显著提升。
Chain-of-Anomaly Thoughts with Large Vision-Language Models
- 设计多智能体推理框架,加入异常导向判断层。
- 低分辨率视频上异常检测F1提升11.8个百分点。
- 适合需要精准识别异常行为的安防场景应用。
基于大视觉语言模型的自动视频监控受限于其对正常状态的固有偏见,常无法检测犯罪行为。尽管思维链推理在语言任务中表现优异,但其推理过程缺乏对异常的归纳偏置,仍倾向于生成正常解释。为此,我们提出异常思维链(Chain-of-Anomaly-Thoughts, CoAT),一种通过最终异常聚焦分类层引入归纳异常偏置的多智能体推理框架。该方法显著提升了异常检测性能,在低分辨率视频上将F1分数提高11.8个百分点;在高分辨率视频中,异常分类准确率提升3.78个百分点。
原文摘要 · Abstract (English)
Automated video surveillance with Large Vision-Language Models is limited by their inherent bias towards normality, often failing to detect crimes. While Chain-of-Thought reasoning strategies show significant potential for improving performance in language tasks, the lack of inductive anomaly biases in their reasoning further steers the models towards normal interpretations. To address this, we propose Chain-of-Anomaly-Thoughts (CoAT), a multi-agent reasoning framework that introduces inductive criminal bias in the reasoning process through a final, anomaly-focused classification layer. Our method significantly improves Anomaly Detection, boosting F1-score by 11.8 p.p. on challenging low-resolution footage and Anomaly Classification by 3.78 p.p. in high-resolution videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。