ALARM通过不确定性量化提升复杂环境下的多模态异常检测可靠性
ALARM: Automated MLLM-Based Anomaly Detection in Complex-EnviRonment Monitoring with Uncertainty Quantification
- 融合不确定性量化与推理链、自我反思等技术
- 在智能家居和伤口图像数据上表现优于现有方法
- 适合需要高可靠决策的工业监控与医疗场景
大型语言模型(LLMs)的发展推动了多模态大语言模型(MLLM)在视觉异常检测(VAD)中的应用,尤其适用于复杂环境。然而,这些环境中异常常具有高度上下文依赖性且模糊,因此不确定性量化(UQ)成为MLLM-VAD系统成功的关键能力。本文提出基于UQ的MLLM-VAD框架ALARM,整合了推理链、自我反思与MLLM集成等质量保障技术,并建立在严格的概率推断与计算流程基础上。通过真实世界智能家居基准数据与伤口图像分类数据的大量实证评估,ALARM展现出优越性能及其在不同领域的通用适用性,支持可靠决策。
原文摘要 · Abstract (English)
The advance of Large Language Models (LLMs) has greatly stimulated research interest in developing multi-modal LLM (MLLM)-based visual anomaly detection (VAD) algorithms that can be deployed in complex environments. The challenge is that in these complex environments, the anomalies are sometimes highly contextual and also ambiguous, and thereby, uncertainty quantification (UQ) is a crucial capacity for an MLLM-based VAD system to succeed. In this paper, we introduce our UQ-supported MLLM-based VAD framework called ALARM. ALARM integrates UQ with quality-assurance techniques like reasoning chain, self-reflection, and MLLM ensemble for robust and accurate performance and is designed based on a rigorous probabilistic inference pipeline and computational process. Extensive empirical evaluations are conducted using the real-world smart-home benchmark data and wound image classification data, which shows ALARM's superior performance and its generic applicability across different domains for reliable decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。