arXiv:2603.15237cs.CV2026-03中稿 · IEEE ICASSP2026被引 2

用物理规则指导对话,让视觉语言模型更准识别异常运动

Multi-turn Physics-informed Vision-language Model for Physics-grounded Anomaly Detection

  • 通过多轮对话注入物体属性与运动规律,引导模型逐步推理
  • 在Phys-AD数据集上达到96.7%的视频级检测AUROC,远超以往最优66.9%
  • 不仅能检出异常,还能给出可解释的因果分析,适合工业质检等场景

视觉语言模型虽具强泛化推理能力,但在需要因果理解动态过程的物理基础异常检测任务中表现受限。现有模型主要依赖外观相关性训练,难以捕捉运动约束,导致对不规则旋转或违反力学规律的动作识别效果差。本文提出一种物理信息指令微调框架,将物体属性、运动范式和动态约束以结构化提示形式注入模型。通过多轮对话逐步传递物理先验,使模型分解因果推理为递进步骤,构建对正常与异常动态的鲁棒表征。在Phys-AD基准测试中,该方法实现96.7%的视频级检测AUROC,显著优于此前最先进水平(66.9%),并获得0.777的LLM评分因果解释质量。研究表明,结构化物理先验可有效提升VLM在动态异常检测中的可靠性。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) demonstrate strong general-purpose reasoning but remain limited in physics-grounded anomaly detection, where causal understanding of dynamics is essential. Existing VLMs, trained predominantly on appearance-centric correlations, fail to capture kinematic constraints, leading to poor performance on anomalies such as irregular rotations or violated mechanical motions. We introduce a physics-informed instruction tuning framework that explicitly encodes object properties, motion paradigms, and dynamic constraints into structured prompts. By delivering these physical priors through multi-turn dialogues, our method decomposes causal reasoning into incremental steps, enabling robust internal representations of normal and abnormal dynamics. Evaluated on the Phys-AD benchmark, our approach achieves 96.7% AUROC in video-level detection--substantially outperforming prior SOTA (66.9%)--and yields superior causal explanations (0.777 LLM score). This work highlights how structured physics priors can transform VLMs into reliable detectors of dynamic anomalies.

异常检测视觉语言模型物理建模因果推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。