arXiv:2412.09819cs.LGcs.SY2024-12被引 10

构建FDM3D打印评测基准,评估大模型在实际制造任务中的表现。

FDM-Bench: A Comprehensive Benchmark for Evaluating Large Language Models in Additive Manufacturing Tasks

  • 设计多层级用户问题与含缺陷G代码的评测数据集。
  • 闭源模型在检测打印异常上优于开源模型,405B版开源模型更擅回答用户问题。
  • 为非专业用户和工程师提供大模型辅助3D打印的可行性参考。

熔融沉积成型(FDM)是一种广泛应用的增材制造技术,因其灵活性和成本效益被广泛用于医疗、航空航天等领域。尽管廉价的FDM设备已普及,但其设计、规划与生产仍需跨学科专业知识,复杂参数管理与打印缺陷修复仍是主要障碍,阻碍了非技术人员参与。大语言模型(LLMs)具备处理文本与代码的能力,有望缓解此问题。然而现有研究多聚焦特定场景,缺乏对多种模型与任务的全面评估。为此,我们提出FDM-Bench,一个面向FDM任务的基准数据集,涵盖不同经验水平用户的查询及包含各类异常的G代码样本。我们评估了GPT-4o、Claude 3.5 Sonnet、Llama-3.1-70B与Llama-3.1-405B四个模型,并由FDM专家对模型输出进行详细评分。结果显示,闭源模型在G-code异常检测上整体表现更优;而Llama-3.1-405B在用户问题响应中略胜一筹。该结果证明FDM-Bench可作为推动大模型在增材制造领域应用的重要基础工具。

原文摘要 · Abstract (English)

Fused Deposition Modeling (FDM) is a widely used additive manufacturing (AM) technique valued for its flexibility and cost-efficiency, with applications in a variety of industries including healthcare and aerospace. Recent developments have made affordable FDM machines accessible and encouraged adoption among diverse users. However, the design, planning, and production process in FDM require specialized interdisciplinary knowledge. Managing the complex parameters and resolving print defects in FDM remain challenging. These technical complexities form the most critical barrier preventing individuals without technical backgrounds and even professional engineers without training in other domains from participating in AM design and manufacturing. Large Language Models (LLMs), with their advanced capabilities in text and code processing, offer the potential for addressing these challenges in FDM. However, existing research on LLM applications in this field is limited, typically focusing on specific use cases without providing comprehensive evaluations across multiple models and tasks. To this end, we introduce FDM-Bench, a benchmark dataset designed to evaluate LLMs on FDM-specific tasks. FDM-Bench enables a thorough assessment by including user queries across various experience levels and G-code samples that represent a range of anomalies. We evaluate two closed-source models (GPT-4o and Claude 3.5 Sonnet) and two open-source models (Llama-3.1-70B and Llama-3.1-405B) on FDM-Bench. A panel of FDM experts assess the models' responses to user queries in detail. Results indicate that closed-source models generally outperform open-source models in G-code anomaly detection, whereas Llama-3.1-405B demonstrates a slight advantage over other models in responding to user queries. These findings underscore FDM-Bench's potential as a foundational tool for advancing research on LLM capabilities in FDM.

3D打印大模型评测FDM工业AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。