arXiv:2509.10282cs.CVcs.LG2025-09被引 3

融合点云、图像与文本,实现无标注3D缺陷检测

MCL-AD: Multimodal Collaboration Learning for Zero-Shot 3D Anomaly Detection

  • 通过多模态提示学习增强跨模态协作能力
  • 在多个数据集上达到当前最佳零样本检测效果
  • 适合数据稀缺场景下的工业缺陷检测应用

零样本3D异常检测旨在不依赖标注训练数据的情况下识别3D物体缺陷,特别适用于数据稀缺、隐私受限或标注成本高的场景。然而,现有方法大多仅关注点云,忽视了互补模态(如RGB图像和文本语义)提供的丰富信息。本文提出MCL-AD框架,通过点云、RGB图像与文本语义的多模态协同学习,实现更优的零样本3D异常检测。具体地,提出多模态提示学习机制(MPLM),通过引入与物体无关的解耦文本提示和多模态对比损失,提升模态内表征能力和模态间协作。此外,设计协同调制机制(CMM),通过联合调制图像引导和点云引导分支,充分挖掘两者互补表征。大量实验表明,所提MCL-AD框架在零样本3D异常检测中达到领先性能。

原文摘要 · Abstract (English)

Zero-shot 3D (ZS-3D) anomaly detection aims to identify defects in 3D objects without relying on labeled training data, making it especially valuable in scenarios constrained by data scarcity, privacy, or high annotation cost. However, most existing methods focus exclusively on point clouds, neglecting the rich semantic cues available from complementary modalities such as RGB images and texts priors. This paper introduces MCL-AD, a novel framework that leverages multimodal collaboration learning across point clouds, RGB images, and texts semantics to achieve superior zero-shot 3D anomaly detection. Specifically, we propose a Multimodal Prompt Learning Mechanism (MPLM) that enhances the intra-modal representation capability and inter-modal collaborative learning by introducing an object-agnostic decoupled text prompt and a multimodal contrastive loss. In addition, a collaborative modulation mechanism (CMM) is proposed to fully leverage the complementary representations of point clouds and RGB images by jointly modulating the RGB image-guided and point cloud-guided branches. Extensive experiments demonstrate that the proposed MCL-AD framework achieves state-of-the-art performance in ZS-3D anomaly detection.

3D检测多模态零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。