arXiv:2601.06415cs.ROcs.AI2026-01中稿 · IEEE SSRR 2025

用视觉语言模型为工业CAD环境添加语义关系,支持机器人仿真与推理。

Semantic Enrichment of CAD-Based Industrial Environments via Scene Graphs for Simulation and Reasoning

  • 基于大视觉语言模型,从CAD文件生成带语义的3D场景图。
  • 在管道结构和功能关系识别上达到高质量的语义标注效果。
  • 适合做工业机器人仿真、智能系统推理的研究者使用。

工业环境中如显示屏、可操作阀门等功能性元件为机器人训练提供了有效路径。在构建需高层场景理解的仿真系统时,环境必须具备同等细节。尽管CAD文件能精确描述几何与视觉信息,但通常缺乏语义、关系和功能信息,限制了仿真与训练能力。本文提出一种离线方法,利用大视觉语言模型(LVLM)构建详细的3D场景图,以补充功能性与可操作元素之间的关系,为动态仿真与推理提供基础。关键成果包括生成语义标签的定量评估结果,以及在管道结构与功能关系识别上的定性表现。所有代码、结果与环境将公开于 https://cad-scenegraph.github.io

原文摘要 · Abstract (English)

Utilizing functional elements in an industrial environment, such as displays and interactive valves, provide effective possibilities for robot training. When preparing simulations for robots or applications that involve high-level scene understanding, the simulation environment must be equally detailed. Although CAD files for such environments deliver an exact description of the geometry and visuals, they usually lack semantic, relational and functional information, thus limiting the simulation and training possibilities. A 3D scene graph can organize semantic, spatial and functional information by enriching the environment through a Large Vision-Language Model (LVLM). In this paper we present an offline approach to creating detailed 3D scene graphs from CAD environments. This will serve as a foundation to include the relations of functional and actionable elements, which then can be used for dynamic simulation and reasoning. Key results of this research include both quantitative results of the generated semantic labels as well as qualitative results of the scene graph, especially in hindsight of pipe structures and identified functional relations. All code, results and the environment will be made available at https://cad-scenegraph.github.io

工业仿真场景图视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。