提出三种多模态对话回复检索方法,端到端效果更优且参数共享提升性能。
On the Effectiveness of Integration Methods for Multimodal Dialogue Response Retrieval
- 分两步与端到端两种融合策略,统一处理文本图像多模态回复检索。
- 端到端方法性能媲美两步法,无需中间步骤,推理更快。
- 参数共享让跨模态任务知识迁移,减少参数量并提升准确率。
多模态聊天机器人已成为对话系统研究与产业中的重点方向。近年来,研究关注对话上下文与回复的多模态特性。本文针对基于检索的多模态对话系统,将回复检索任务形式化为三个子任务的组合。提出基于两步法和端到端法的三种集成方法,并对比其优劣。在两个数据集上的实验表明,端到端方法在无需中间步骤的情况下达到与两步法相当的性能。此外,参数共享策略不仅减少参数数量,还通过跨子任务与跨模态的知识迁移提升了整体表现。
原文摘要 · Abstract (English)
Multimodal chatbots have become one of the major topics for dialogue systems in both research community and industry. Recently, researchers have shed light on the multimodality of responses as well as dialogue contexts. This work explores how a dialogue system can output responses in various modalities such as text and image. To this end, we first formulate a multimodal dialogue response retrieval task for retrieval-based systems as the combination of three subtasks. We then propose three integration methods based on a two-step approach and an end-to-end approach, and compare the merits and demerits of each method. Experimental results on two datasets demonstrate that the end-to-end approach achieves comparable performance without an intermediate step in the two-step approach. In addition, a parameter sharing strategy not only reduces the number of parameters but also boosts performance by transferring knowledge across the subtasks and the modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。