首个面向智能家居的视频异常检测基准,评估多模态大模型表现。
SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Large Language Models
- 构建涵盖7类异常的1203段智能家居视频数据集。
- 现有大模型检测准确率有限,新框架提升11.62%精度。
- 适合研究智能安防、多模态大模型应用的学者使用。
视频异常检测(VAD)对提升安全与保障至关重要,能识别各类环境中的异常事件。然而,现有基准多针对通用场景,忽视了智能家居的独特需求。为此,我们提出SmartHome-Bench,首个专为智能家居场景设计的综合性视频异常检测基准,聚焦多模态大语言模型(MLLMs)的评估能力。该基准包含1,203段由智能家庭摄像头录制的视频,依据全新异常分类体系划分为七类,如野生动物、老人照护、婴儿看护等。每段视频均配有异常标签、详细描述及推理过程。我们进一步研究了MLLM在VAD中的适配方法,评估了多种开源与闭源模型在不同提示策略下的表现。结果揭示当前模型在异常检测上存在显著局限。为此,我们提出一种基于分类驱动的反思式大模型链(TRLC),在检测准确率上实现11.62%的显著提升。数据集与代码已公开于https://github.com/Xinyi-0724/SmartHome-Bench-LLM。
原文摘要 · Abstract (English)
Video anomaly detection (VAD) is essential for enhancing safety and security by identifying unusual events across different environments. Existing VAD benchmarks, however, are primarily designed for general-purpose scenarios, neglecting the specific characteristics of smart home applications. To bridge this gap, we introduce SmartHome-Bench, the first comprehensive benchmark specially designed for evaluating VAD in smart home scenarios, focusing on the capabilities of multi-modal large language models (MLLMs). Our newly proposed benchmark consists of 1,203 videos recorded by smart home cameras, organized according to a novel anomaly taxonomy that includes seven categories, such as Wildlife, Senior Care, and Baby Monitoring. Each video is meticulously annotated with anomaly tags, detailed descriptions, and reasoning. We further investigate adaptation methods for MLLMs in VAD, assessing state-of-the-art closed-source and open-source models with various prompting techniques. Results reveal significant limitations in the current models' ability to detect video anomalies accurately. To address these limitations, we introduce the Taxonomy-Driven Reflective LLM Chain (TRLC), a new LLM chaining framework that achieves a notable 11.62% improvement in detection accuracy. The benchmark dataset and code are publicly available at https://github.com/Xinyi-0724/SmartHome-Bench-LLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。