arXiv:2603.23122cs.CV2026-03

让机器人主动调整视角和消除干扰,提升工业视觉异常检测鲁棒性

PiCo: Active Manifold Canonicalization for Robust Robotic Visual Anomaly Detection

  • 通过主动重定向物体和多级去噪,将观测投影到不变的语义流形上
  • 在静态场景下达93.7% O-AUROC,闭环控制下准确率达98.5%
  • 适合需要高鲁棒性的工业机器人视觉系统部署

工业机器人视觉异常检测受6-DoF姿态变化及光照、阴影等不稳定条件制约,内在语义异常与物理扰动共存且相互作用。为突破限制,提出从被动特征学习到主动流形归一化的范式转变。本文提出统一框架PiCo(Pose-in-Condition Canonicalization),主动将观测投影至条件不变的规范流形。其采用级联机制:第一阶段主动物理归一化,通过重定向物体减少几何不确定性;第二阶段神经潜在归一化,包含输入层的辐射校正、特征层的潜在精炼与语义层的上下文推理,逐级消除多尺度干扰因素。在大规模M2AD基准测试中,PiCo实现93.7% O-AUROC(较此前方法提升3.7%),在主动闭环场景下达到98.5%准确率,验证了主动流形归一化对鲁棒具身感知的关键作用。

原文摘要 · Abstract (English)

Industrial deployment of robotic visual anomaly detection (VAD) is fundamentally constrained by passive perception under diverse 6-DoF pose configurations and unstable operating conditions such as illumination changes and shadows, where intrinsic semantic anomalies and physical disturbances coexist and interact. To overcome these limitations, a paradigm shift from passive feature learning to Active Canonicalization is proposed. PiCo (Pose-in-Condition Canonicalization) is introduced as a unified framework that actively projects observations onto a condition-invariant canonical manifold. PiCo operates through a cascaded mechanism. The first stage, Active Physical Canonicalization, enables a robotic agent to reorient objects in order to reduce geometric uncertainty at its source. The second stage, Neural Latent Canonicalization, adopts a three-stage denoising hierarchy consisting of photometric processing at the input level, latent refinement at the feature level, and contextual reasoning at the semantic level, progressively eliminating nuisance factors across representational scales. Extensive evaluations on the large-scale M2AD benchmark demonstrate the superiority of this paradigm. PiCo achieves a state-of-the-art 93.7% O-AUROC, representing a 3.7% improvement over prior methods in static settings, and attains 98.5% accuracy in active closed-loop scenarios. These results demonstrate that active manifold canonicalization is critical for robust embodied perception.

视觉异常检测机器人感知流形学习主动感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。