arXiv:2508.08652cs.AI2025-08被引 1

用大模型自动检查模拟训练中的沟通合规性,无需额外训练。

Prompt-and-Check: Using Large Language Models to Evaluate Communication Protocol Compliance in Simulation-Based Training

  • 用上下文提示让大模型分析语音转录内容判断是否合规。
  • 在海事模拟中准确识别90%以上的合规项,接近专家水平。
  • 可在普通显卡上运行,适合培训系统快速部署。

在安全关键领域,程序化沟通的准确评估对模拟训练至关重要,遵守检查清单反映实际操作能力。本文提出一种轻量级、可部署的方法,利用开源大语言模型(LLMs)进行基于提示的推理,能在消费级GPU上高效运行。我们提出Prompt-and-Check方法,通过富含上下文的提示,仅根据语音转录内容判断协议中每个检查项是否完成。在海事领域开展案例研究,参与者执行相同模拟任务,实验使用LLama 2 7B、LLaMA 3 8B和Mistral 7B模型,在RTX 4070 GPU上本地运行。针对每项检查,将相关语料片段输入模型,输出合规判断。通过分类准确率和一致度评分,对比模型输出与专家标注的真值。结果表明,提示工程可实现无需任务特训的有效上下文推理。该研究凸显了大模型在增强复盘、绩效反馈与自动化评估中的实用价值。

原文摘要 · Abstract (English)

Accurate evaluation of procedural communication compliance is essential in simulation-based training, particularly in safety-critical domains where adherence to compliance checklists reflects operational competence. This paper explores a lightweight, deployable approach using prompt-based inference with open-source large language models (LLMs) that can run efficiently on consumer-grade GPUs. We present Prompt-and-Check, a method that uses context-rich prompts to evaluate whether each checklist item in a protocol has been fulfilled, solely based on transcribed verbal exchanges. We perform a case study in the maritime domain with participants performing an identical simulation task, and experiment with models such as LLama 2 7B, LLaMA 3 8B and Mistral 7B, running locally on an RTX 4070 GPU. For each checklist item, a prompt incorporating relevant transcript excerpts is fed into the model, which outputs a compliance judgment. We assess model outputs against expert-annotated ground truth using classification accuracy and agreement scores. Our findings demonstrate that prompting enables effective context-aware reasoning without task-specific training. This study highlights the practical utility of LLMs in augmenting debriefing, performance feedback, and automated assessment in training environments.

大模型应用模拟训练合规检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。