用大模型优化工业机器人对话数据,提升人机交互准确率
IRWOZ 2.0: A Large Language Model-driven Dialogue Dataset for Industrial Robot Conversations

- 用Mistral/Claude-3.5生成并修正对话,降低噪声
- 4个工业场景共390组对话,状态追踪准确率显著提升
- 适合研究工业人机对话的学者与工程师使用
IRWOZ通过领域特定标注提升了工业人机交互对话系统的性能,但其初版存在大量对话状态和语句噪声,影响状态追踪准确性。本文提出IRWOZ 2.0,利用大语言模型(Mistral/Claude-3.5)增强生成与质量优化,涵盖装配、配送、定位、重定位四个工业场景,共390组对话,经人工校正与自动错别字修复。在对话状态追踪基准测试中,GPT-2的BLEU-4得分从0.1651提升至0.5604。为支持工业人机交互研究,数据集已公开发布于https://ieee-dataport.org/documents/irwoz-20-large-language-model-driven-dialogue-dataset-industrial-robot-conversations。
原文摘要 · Abstract (English)
IRWOZ has improved industrial human-robot interaction (HRI) dialogue systems through domain-specific annotations. However, its initial version contains substantial noise in dialogue states and utterances, limiting state-tracking accuracy. We introduce IRWOZ 2.0, which addresses these limitations through large language model (LLM) enhanced generation (Mistral/Claude-3.5) and quality refinements. Our improved dataset expands to 390 dialogues across 4 industrial domains (Assembly, Delivery, Position, Relocation), featuring manual corrections and automated typo removal. Benchmark experiments on dialogue state tracking demonstrate significant improvements, with GPT-2's BLEU-4 score increasing from 0.1651 to 0.5604 compared to original IRWOZ. To support industrial HRI research, we publicly released IRWOZ 2.0 dataset at https://ieee-dataport.org/documents/irwoz-20-large-language-model-driven-dialogue-dataset-industrial-robot-conversations
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。