arXiv:2602.20551cs.CV2026-02

用CAD模型生成的几何提示,实现工业零件的精准分割。

CAD-Prompted SAM3: Geometry-Conditioned Instance Segmentation for Industrial Objects

  • 以CAD模型多视角渲染图为几何提示,替代语言或外观描述。
  • 在合成数据上训练,实现单阶段、无需微调的分割预测。
  • 适合材质颜色多变但形状固定的工业零件,如3D打印件。

语言提示在描述制造与3D打印环境中罕见、实例特异或难以描述的物体时存在表达力不足的问题。图像示例虽可作为替代,但主要依赖颜色、纹理等外观特征,而这些特征常与零件的几何身份无关。工业场景中,同一零件可能使用不同材料、表面处理或颜色,使基于外观的提示不可靠。相比之下,这些物体通常由精确的CAD模型定义,能完整刻画其标准几何结构。本文提出基于SAM3的CAD提示分割框架,采用CAD模型的多视角渲染图作为提示输入,实现与表面外观无关的几何条件化分割。模型通过在仿真环境中生成的多种视角和场景上下文的网格渲染合成数据进行训练。该方法支持单阶段、无需微调的掩码预测,将可提示分割扩展至无法仅凭语言或外观可靠描述的对象。

原文摘要 · Abstract (English)

Verbal-prompted segmentation is inherently limited by the expressiveness of natural language and struggles with uncommon, instance-specific, or difficult-to-describe objects: scenarios frequently encountered in manufacturing and 3D printing environments. While image exemplars provide an alternative, they primarily encode appearance cues such as color and texture, which are often unrelated to a part's geometric identity. In industrial settings, a single component may be produced in different materials, finishes, or colors, making appearance-based prompting unreliable. In contrast, such objects are typically defined by precise CAD models that capture their canonical geometry. We propose a CAD-prompted segmentation framework built on SAM3 that uses canonical multi-view renderings of a CAD model as prompt input. The rendered views provide geometry-based conditioning independent of surface appearance. The model is trained using synthetic data generated from mesh renderings in simulation under diverse viewpoints and scene contexts. Our approach enables single-stage, CAD-prompted mask prediction, extending promptable segmentation to objects that cannot be robustly described by language or appearance alone.

实例分割工业视觉几何提示CAD

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。