用物理运动预测未来语义,提前发现视频异常。
Latent Clarity: Bridging World-Model Kinematics to Semantic Manifolds for Video Anomaly Anticipation
- 构建可预测的潜在空间,将物理运动映射到语义空间
- 在UCF-Crime上达0.8994的异常检测准确率,零样本迁移有效
- 首次证明未来预测比当前观测更具语义可分性,适合异常预警研究
连续视频异常检测长期依赖反应式多实例学习(MIL),将时空特征压缩为标量分数。本文提出PULS(预测统一潜在空间)框架,包含两个模块:490M参数的KSD桥接器(从物理张量映射至语义空间)和16.8M参数的前瞻状态预测器(ASP)。KSD桥接器将V-JEPA 2物理张量映射至2048维的Qwen3-VL-Embedding-2B文本对齐超球面,仅在UCF-Crime子集上训练,即实现块级AUROC 0.8994(UCF-Crime)和0.8162(出域XD-Violence),无需MIL或层级融合。提出并验证潜空间清晰度假说:因JEPA的时间预测器剔除随机像素噪声而保留运动学信息,未来表征比当前观测更易语义分离。ASP进一步锐化未来潜在表示,在14类零样本视频问答任务中达44.5%平均准确率(优于观测基线9.6个百分点)。将ASP应用于观测张量时准确率降至7.3%(随机水平),证明前瞻与观测处于不同子流形。采用三轨提前量协议与L1意外门控,于T-0.5秒处获得+8.9个百分点的前瞻优势(p < 0.001,N = 1,000置换检验),分离物理前瞻与静态场景先验。零样本迁移至XD-Violence证实牛顿不变运动表示具备强泛化能力。
原文摘要 · Abstract (English)
Continuous video anomaly detection is dominated by reactive Multiple Instance Learning (MIL) that collapses spatiotemporal features into scalar scores. We introduce PULS (Predictive Unified Latent Space), a continuous semantic world-model pipeline comprising two modules: a 490M-parameter KSD Bridge (Kinematic-to-Semantic Distillation) and a 16.8M-parameter Anticipatory State Predictor (ASP). The KSD Bridge maps V-JEPA 2 physical tensors into the 2048-d Qwen3-VL-Embedding-2B text-aligned hypersphere, trained on a subset of UCF-Crime. This translation alone yields a chunk-level AUROC of 0.8994 for UCF-Crime and 0.8162 for out-of-distribution XD-Violence without MIL or hierarchical fusion. We introduce and validate the Latent Clarity Hypothesis: because JEPA's temporal predictor discards aleatoric pixel noise while preserving kinematics, anticipated future representations are more semantically separable than observed presents. The ASP sharpens these anticipated future latents, achieving 44.5% mean 14-way zero-shot VQA accuracy (exceeding observation baseline by +9.6 pp). Applying the ASP to Observation Tensors collapses accuracy to 7.3% (random chance), proving Anticipation and Observation occupy distinct sub-manifolds. A Triple-Track Lead-Time protocol with an L1-surprise gate yields a peak +8.9 pp anticipatory advantage at T-0.5s (p < 0.001, N = 1,000 permutation), separating physical anticipation from static scene priors. Zero-shot transfer to XD-Violence confirms that Newtonian-invariant kinematic representations generalize out-of-distribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。