测试大模型对焊接质量的判断能力,发现其在真实场景中仍有潜力。
Do Multimodal Large Language Models Understand Welding?
- 用真实焊接图像测试大模型,结合专家标注数据评估性能。
- 模型在在线图片上表现更好,但对真实焊接图像也能准确判断。
- 提出WeldPrompt提示策略,提升推理可靠性,但效果因场景而异。
本文研究多模态大模型(MLLMs)在高技能生产任务中的表现,聚焦焊接质量评估。基于由领域专家标注的真实与网络焊接图像数据集,评估两种先进MLLMs在房车/船舶、航空和农业三个场景下的表现。尽管模型在在线图像上表现更优,可能源于前期训练或记忆,但在未见过的真实焊接图像上也表现出较好性能。为此,本文提出WeldPrompt,结合思维链生成与上下文学习的提示策略,以减少幻觉并增强推理能力。该策略在部分场景下提升了模型召回率,但跨场景表现不一致。结果表明,当前MLLMs在高风险技术领域仍具潜力但存在局限,强调微调、领域数据和复杂提示策略对提升可靠性的关键作用。研究为工业多模态学习开辟新方向。
原文摘要 · Abstract (English)
This paper examines the performance of Multimodal LLMs (MLLMs) in skilled production work, with a focus on welding. Using a novel data set of real-world and online weld images, annotated by a domain expert, we evaluate the performance of two state-of-the-art MLLMs in assessing weld acceptability across three contexts: RV \& Marine, Aeronautical, and Farming. While both models perform better on online images, likely due to prior exposure or memorization, they also perform relatively well on unseen, real-world weld images. Additionally, we introduce WeldPrompt, a prompting strategy that combines Chain-of-Thought generation with in-context learning to mitigate hallucinations and improve reasoning. WeldPrompt improves model recall in certain contexts but exhibits inconsistent performance across others. These results underscore the limitations and potentials of MLLMs in high-stakes technical domains and highlight the importance of fine-tuning, domain-specific data, and more sophisticated prompting strategies to improve model reliability. The study opens avenues for further research into multimodal learning in industry applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。