arXiv:2409.14247cs.CLcs.HC2024-09EMNLP被引 6

构建多模态纠错数据集,测试大模型在对话中修复误解的能力

Repairs in a Block World: A New Benchmark for Handling User Corrections with Multi-Modal Language Models

  • 构建包含指代歧义的多模态纠错数据集BlockWorld-Repairs
  • 现有模型在纠错任务上表现远低于人类水平
  • 针对性微调损失可提升模型对交互纠错的学习能力

对话中,听者可能初始误解说话者并做出错误回应,常引发说话者在下一轮以第三位置修正(TPR)方式进行纠正。处理此类纠错序列的能力对对话系统至关重要。本文首次收集、分析并公开发布BlockWorld-Repairs:一个面向指令执行操作任务的多模态TPR序列数据集,该任务设计上充满指代歧义。我们利用此数据集评估多个前沿视觉语言模型(VLM)在多种设置下的表现,重点考察其处理并准确响应TPR以恢复误沟通的能力。结果表明,与人类相比,所有模型在此任务上均显著表现不足。进一步实验显示,通过在微调中引入针对相关标记的专用损失,可有效提升模型性能并增强对新场景的泛化能力。研究结果表明,当前模型尚未具备在常见需纠错的多模态协作场景中部署的能力,强调需设计更优的训练策略与目标以促进从交互中学习。代码与数据已开源。

原文摘要 · Abstract (English)

In dialogue, the addressee may initially misunderstand the speaker and respond erroneously, often prompting the speaker to correct the misunderstanding in the next turn with a Third Position Repair (TPR). The ability to process and respond appropriately to such repair sequences is thus crucial in conversational AI systems. In this paper, we first collect, analyse, and publicly release BlockWorld-Repairs: a dataset of multi-modal TPR sequences in an instruction-following manipulation task that is, by design, rife with referential ambiguity. We employ this dataset to evaluate several state-of-the-art Vision and Language Models (VLM) across multiple settings, focusing on their capability to process and accurately respond to TPRs and thus recover from miscommunication. We find that, compared to humans, all models significantly underperform in this task. We then show that VLMs can benefit from specialised losses targeting relevant tokens during fine-tuning, achieving better performance and generalising better to new scenarios. Our results suggest that these models are not yet ready to be deployed in multi-modal collaborative settings where repairs are common, and highlight the need to design training regimes and objectives that facilitate learning from interaction. Our code and data are available at www.github.com/JChiyah/blockworld-repairs

多模态对话系统纠错视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。