arXiv:2512.20387cs.AIcs.CL2025-12

用视觉和语言生成可执行的工业仿真代码,实现数字孪生的智能构建。

Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems

  • 通过图文输入自动生成可运行的FlexScript代码,融合视觉与语言理解。
  • 在12万组图文代码数据上训练,结构准确率接近完美,执行成功率高。
  • 专为工业仿真设计新评估指标,适合智能制造与自动化研究者参考。

我们提出一种视觉-语言仿真模型(VLSM),通过统一视觉与文本理解,从布局草图和自然语言提示中合成可执行的FlexScript代码,实现工业仿真系统的跨模态推理。为此,研究构建了首个大规模生成式数字孪生数据集,包含超过12万组提示-草图-代码三元组,支持文本描述、空间结构与仿真逻辑间的多模态学习。同时,提出三种新评估指标:结构有效性率(SVR)、参数匹配率(PMR)和执行成功率(ESR),全面评估结构完整性、参数准确性与仿真器可执行性。通过系统性消融实验,验证不同视觉编码器、连接器与代码预训练语言模型的效果,所提模型达到近乎完美的结构准确率与高执行鲁棒性。本工作为融合视觉推理与语言理解的可执行工业仿真系统奠定了基础。

原文摘要 · Abstract (English)

We propose a Vision-Language Simulation Model (VLSM) that unifies visual and textual understanding to synthesize executable FlexScript from layout sketches and natural-language prompts, enabling cross-modal reasoning for industrial simulation systems. To support this new paradigm, the study constructs the first large-scale dataset for generative digital twins, comprising over 120,000 prompt-sketch-code triplets that enable multimodal learning between textual descriptions, spatial structures, and simulation logic. In parallel, three novel evaluation metrics, Structural Validity Rate (SVR), Parameter Match Rate (PMR), and Execution Success Rate (ESR), are proposed specifically for this task to comprehensively evaluate structural integrity, parameter fidelity, and simulator executability. Through systematic ablation across vision encoders, connectors, and code-pretrained language backbones, the proposed models achieve near-perfect structural accuracy and high execution robustness. This work establishes a foundation for generative digital twins that integrate visual reasoning and language understanding into executable industrial simulation systems. Project page: https://danielhsu2014.github.io/GDT-VLSM-project/

数字孪生生成模型工业仿真多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。