通过大小模型协作,提升电商少样本多模态对话意图识别效果
Knowledge-Decoupled Synergetic Learning: An MLLM based Collaborative Approach to Few-shot Multimodal Dialogue Intention Recognition
- 用小模型提取可解释规则,大模型后训练,解耦知识干扰
- 在淘宝真实数据上,加权F1提升6.37%和6.28%
- 适合需要少样本多模态理解的电商对话系统开发者
少样本多模态对话意图识别是电子商务领域的重要挑战。以往方法主要通过后训练提升模型分类能力,但我们的分析发现,该任务涉及两个相互关联的任务,导致多任务学习中出现此消彼长现象,根源在于训练过程中权重矩阵更新叠加引发的知识干扰。为此,我们提出知识解耦协同学习(KDSL)框架,利用小型模型将知识转化为可解释规则,同时对大型模型进行后训练。通过大、小多模态大语言模型协同预测,显著提升性能。在两个真实淘宝数据集上,相比当前最优方法,线上加权F1分别提升6.37%和6.28%,验证了该框架的有效性。
原文摘要 · Abstract (English)
Few-shot multimodal dialogue intention recognition is a critical challenge in the e-commerce domainn. Previous methods have primarily enhanced model classification capabilities through post-training techniques. However, our analysis reveals that training for few-shot multimodal dialogue intention recognition involves two interconnected tasks, leading to a seesaw effect in multi-task learning. This phenomenon is attributed to knowledge interference stemming from the superposition of weight matrix updates during the training process. To address these challenges, we propose Knowledge-Decoupled Synergetic Learning (KDSL), which mitigates these issues by utilizing smaller models to transform knowledge into interpretable rules, while applying the post-training of larger models. By facilitating collaboration between the large and small multimodal large language models for prediction, our approach demonstrates significant improvements. Notably, we achieve outstanding results on two real Taobao datasets, with enhancements of 6.37\% and 6.28\% in online weighted F1 scores compared to the state-of-the-art method, thereby validating the efficacy of our framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。