用大模型增强制造任务指导中的问答系统,提升技术员操作支持效率
QA-TOOLBOX: Conversational Question-Answering for process task guidance in manufacturing
- 基于制造场景的20万+问答对,结合文档与视频实现多模态数据增强
- 在无参考条件下评估多个开源大模型,发现其对流程理解存在显著差距
- 通过专家验证和众包评分,建立可信的性能评估框架,适合工业AI研究者
本文探索利用大语言模型(LLMs)为先进制造环境中的任务指导系统进行数据增强。数据集包含20万+条来自技术人员交互的问答对,涵盖流程规范文档、动作与物体的时序序列信息,并基于叙述和/或视频演示进行标注。研究旨在分析该任务的复杂性,评估现有LLMs在任务支持中的表现。通过为每个主流开源LLM构建基线,在无参考设置下使用大模型作为评判者(LLM-as-a-judge),并与众包工作者评分结果对比,同时由专家验证评分一致性。结果表明,任务理解需整合文档、动作时序与上下文线索,当前大模型在此类任务中仍存在明显局限。
原文摘要 · Abstract (English)
In this work we explore utilizing LLMs for data augmentation for manufacturing task guidance system. The dataset consists of representative samples of interactions with technicians working in an advanced manufacturing setting. The purpose of this work to explore the task, data augmentation for the supported tasks and evaluating the performance of the existing LLMs. We observe that that task is complex requiring understanding from procedure specification documents, actions and objects sequenced temporally. The dataset consists of 200,000+ question/answer pairs that refer to the spec document and are grounded in narrations and/or video demonstrations. We compared the performance of several popular open-sourced LLMs by developing a baseline using each LLM and then compared the responses in a reference-free setting using LLM-as-a-judge and compared the ratings with crowd-workers whilst validating the ratings with experts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。