arXiv:2410.23988cs.CV2024-10被引 2

JEMA通过多模态对齐实现激光金属沉积过程的高效监控与泛化预测。

JEMA: A Joint Embedding Framework for Scalable Co-Learning with Multimodal Alignment

  • 融合多视角图像与工艺参数,用对比学习构建可迁移的语义表示。
  • 在多模态下性能提升8%,单模态下提升1%,且无需大量微调。
  • 适合工业级金属增材制造中数据少、解释性要求高的场景。

本文提出JEMA(联合嵌入与多模态对齐),一种面向激光金属沉积(LMD)的协同学习框架。随着工业5.0发展,高效过程监控日益重要,但数据有限与AI黑箱特性制约其应用。JEMA利用多视图图像与工艺参数等多模态数据,学习可迁移的语义表征。通过监督对比损失函数,实现仅用主模态即可鲁棒训练,降低硬件与计算开销。实验验证其在熔池几何预测等下游任务中的泛化能力,无需大量微调。结果表明,在多模态设置下性能提升8%,单模态下提升1%,优于传统监督对比学习。所学嵌入还能预测元数据,增强可解释性并评估元数据贡献。该框架为集成多传感器数据与元数据提供基础,支持LMD及更广泛领域的多样化下游任务。

原文摘要 · Abstract (English)

This work introduces JEMA (Joint Embedding with Multimodal Alignment), a novel co-learning framework tailored for laser metal deposition (LMD), a pivotal process in metal additive manufacturing. As Industry 5.0 gains traction in industrial applications, efficient process monitoring becomes increasingly crucial. However, limited data and the opaque nature of AI present challenges for its application in an industrial setting. JEMA addresses this challenges by leveraging multimodal data, including multi-view images and metadata such as process parameters, to learn transferable semantic representations. By applying a supervised contrastive loss function, JEMA enables robust learning and subsequent process monitoring using only the primary modality, simplifying hardware requirements and computational overhead. We investigate the effectiveness of JEMA in LMD process monitoring, focusing specifically on its generalization to downstream tasks such as melt pool geometry prediction, achieved without extensive fine-tuning. Our empirical evaluation demonstrates the high scalability and performance of JEMA, particularly when combined with Vision Transformer models. We report an 8% increase in performance in multimodal settings and a 1% improvement in unimodal settings compared to supervised contrastive learning. Additionally, the learned embedding representation enables the prediction of metadata, enhancing interpretability and making possible the assessment of the added metadata's contributions. Our framework lays the foundation for integrating multisensor data with metadata, enabling diverse downstream tasks within the LMD domain and beyond.

多模态学习金属增材制造对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。