arXiv:2605.24503cs.CVcs.AI2026-05

构建食品监管场景的可解释合规分析基准,助力厨房安全智能监控。

FoodMonitor: Benchmarking MLLMs for Explainable Compliance Analysis

论文配图:FoodMonitor: Benchmarking MLLMs for Explainable Compliance Analysis
图 1 · 摘自论文原文
  • 设计双通道视频数据集,覆盖人员与环境违规行为
  • 最佳模型仅得0.360分,空间定位与规则理解成主要瓶颈
  • 提供诊断性错误分析,指导未来模型改进方向

随着AI驱动的合规监控在公共治理与工业安全中日益重要,提供可验证证据与可追溯问责信号至关重要。然而,现有视频异常检测数据集仅关注事件级二分类,缺乏真实合规场景所需的规则驱动、可解释分析能力。我们提出FoodMonitor,一个用于商业厨房监控可解释合规分析的基准。该数据集包含477段视频片段,共3,307个违规标注,采用双通道设计,涵盖人员与环境层面违规。每个标注明确指出违反的规则、具体不合规行为及责任人,并附帧级边界框。我们建立统一评估协议,采用两阶段匹配机制分别评估空间定位与语义理解,并引入综合指标$C_{\text{score}}$平衡环境与人员检测性能。对多个先进多模态大模型的系统评估显示,最优模型仅达0.360 $C_{\text{score}}$,空间定位与细粒度规则理解成为主要瓶颈。分析识别出两类典型失败模式:以定位为主与以语义为主的错误,为未来模型发展提供诊断依据。

原文摘要 · Abstract (English)

As AI-powered compliance monitoring becomes increasingly important in public governance and industrial safety, the ability to provide verifiable evidence and traceable accountability signals is essential. However, existing video anomaly detection datasets focus on event-level binary classification, lacking the rule-driven, explainable analysis required for real-world compliance scenarios. We introduce FoodMonitor, a benchmark for explainable compliance analysis in commercial kitchen surveillance. FoodMonitor comprises 477 video clips with 3,307 violation annotations across a dual-channel design covering both person-level and environment-level violations. Each annotation specifies which rule was violated, what non-compliant behavior occurred, and who committed it with frame-level bounding boxes. We establish a unified evaluation protocol with a two-stage matching mechanism that separately assesses spatial localization and semantic understanding, along with a composite metric ($C_{\text{score}}$) that balances environment and person detection performance. Systematic evaluation of several state-of-the-art multimodal large language models reveals that the best-performing model achieves only 0.360 $C_{\text{score}}$, with spatial localization and fine-grained rule understanding emerging as the primary bottlenecks. Our analysis identifies two distinct failure modes: localization-dominated errors and semantics-dominated errors, providing diagnostic insights for future model development.

合规分析多模态模型视频理解可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。