arXiv:2606.21579cs.CVcs.AI2026-06

用预训练视频语言模型实现零样本流程错误检测,效果媲美有监督方法。

The Unreasonable Effectiveness of VLMs for Zero-shot Procedural Mistake Detection

论文配图:The Unreasonable Effectiveness of VLMs for Zero-shot Procedural Mistake Detection
图 1 · 摘自论文原文
  • 仅用一个预训练视频语言模型,联合完成错误检测与动作分割。
  • 在EgoPER和CaptainCook4D上性能接近甚至超过有监督方法。
  • 无需特定任务训练,适合跨领域快速部署的场景。

流程错误检测在质量控制与用户辅助中至关重要。现有方法依赖多阶段流水线,需针对任务训练数据,限制了通用性。为此,我们提出零样本流程错误检测框架ZeProM,仅使用单一预训练视频语言模型(VLM),联合解决流程错误检测与时间动作分割。在两个基准数据集EgoPER和CaptainCook4D上的评估显示,该框架可成功执行任务,性能接近甚至超越完全有监督方法。例如,在所有五个EgoPER任务上,平均提升EDA 4.4点,[email protected]提升2.0点。结果表明,统一方法具有巨大潜力,有望推动领域从复杂流水线转向更通用的解决方案。

原文摘要 · Abstract (English)

Procedural mistake detection is important for quality control and user assistance across many disciplines. Recent work in this field has achieved significant gains by using the reasoning capabilities of Video-Language Models (VLMs) as components within multi-stage pipelines, which consist of separate modules for supervised temporal action segmentation, error detection, and explainability. Consequently, they remain dependent on tailored training datasets and require task-specific training, limiting their wider applicability. To remedy this, we introduce zero-shot procedural mistake detection and propose a unified Zero-shot Procedural Mistake detection (ZeProM) framework that jointly solves procedural mistake detection and temporal action segmentation with a single pre-trained VLM. By evaluating our framework on two canonical mistake detection benchmarks, EgoPER and CaptainCook4D, we find that ZeProM can perform these tasks successfully, while approaching, or even outperforming, the performance of fully supervised methods. For instance, we achieve a 4.4 point improvement in EDA and a 2.0 point improvement in [email protected] on average over all five EgoPER tasks compared to the strongest supervised methods. Overall, our results show the potential of unified methods for procedural mistake detection, and we hope this will steer the field away from highly complex pipelines and toward more generally applicable solutions.

零样本视频语言模型流程检测统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。