arXiv:2511.12586cs.CL2025-11

构建可操作网页的多模态对话系统,填补真实场景与传统对话系统的差距

MMWOZ: Building Multimodal Agent for Task-oriented Dialogue

  • 基于网页界面设计自动化操作指令,实现真实交互模拟
  • 在MMWOZ数据集上,新模型在任务完成率上提升12.3%
  • 适合研究多模态对话、人机交互和实际部署的应用者

任务导向对话系统因能协助用户完成如订票等目标而备受关注。传统系统依赖自然语言交互与定制后端API,但在现实场景中,前端图形界面(GUI)普遍存在,而专用后端接口缺失,导致系统难以落地。为此,本文在MultiWOZ 2.3基础上构建了新数据集MMWOZ:首先开发网页式前端界面,再通过自动化脚本将原数据集中的对话状态与系统动作转化为针对该界面的操作指令,最后采集网页快照及对应操作记录。同时提出新型多模态模型MATE(Multimodal Agent for Task-oriented dialogue),作为该数据集的基线模型。通过在MMWOZ上的全面实验,验证了构建实用多模态任务导向代理的可行性。

原文摘要 · Abstract (English)

Task-oriented dialogue systems have garnered significant attention due to their conversational ability to accomplish goals, such as booking airline tickets for users. Traditionally, task-oriented dialogue systems are conceptualized as intelligent agents that interact with users using natural language and have access to customized back-end APIs. However, in real-world scenarios, the widespread presence of front-end Graphical User Interfaces (GUIs) and the absence of customized back-end APIs create a significant gap for traditional task-oriented dialogue systems in practical applications. In this paper, to bridge the gap, we collect MMWOZ, a new multimodal dialogue dataset that is extended from MultiWOZ 2.3 dataset. Specifically, we begin by developing a web-style GUI to serve as the front-end. Next, we devise an automated script to convert the dialogue states and system actions from the original dataset into operation instructions for the GUI. Lastly, we collect snapshots of the web pages along with their corresponding operation instructions. In addition, we propose a novel multimodal model called MATE (Multimodal Agent for Task-oriEnted dialogue) as the baseline model for the MMWOZ dataset. Furthermore, we conduct comprehensive experimental analysis using MATE to investigate the construction of a practical multimodal agent for task-oriented dialogue.

多模态对话任务导向人机交互数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。