轻量模型也能生成有逻辑的多模态情感推理链并准确分类。
Resource-Limited Joint Multimodal Sentiment Reasoning and Classification via Chain-of-Thought Enhancement and Distillation
- 用师生蒸馏框架让小模型模仿大模型生成情感推理链。
- 30亿参数模型在4个数据集上表现接近大模型,且推理过程可解释。
- 适合移动端、嵌入式等算力受限场景下的多模态情感分析。
社交媒体上丰富的多模态内容推动了多模态情感分析(MSA)的发展,大型语言模型(LLMs)进一步加速了该领域的进展。现有方法主要依赖参数庞大的多模态大模型进行情感分类,忽视了资源受限环境下自主生成情感推理链的能力。为此,本文聚焦于资源受限的联合多模态情感推理与分类任务(JMSRC),仅使用轻量级模型同时完成情感推理链生成与分类。提出多模态思维链推理蒸馏模型MulCoT-RD,采用“教师-助教-学生”蒸馏范式,以应对部署约束。首先利用高性能多模态大模型(MLLM)生成初始推理数据,并训练一个中等规模助教模型;随后联合训练一个轻量级学生模型,实现高效的多模态情感推理生成与分类。在四个数据集上的大量实验表明,仅30亿参数的MulCoT-RD在JMSRC任务上表现强劲,具备良好的泛化能力与增强的可解释性。
原文摘要 · Abstract (English)
The surge in rich multimodal content on social media platforms has greatly advanced Multimodal Sentiment Analysis (MSA), with Large Language Models (LLMs) further accelerating progress in this field. Current approaches primarily leverage the knowledge and reasoning capabilities of parameter-heavy (Multimodal) LLMs for sentiment classification, overlooking autonomous multimodal sentiment reasoning generation in resource-constrained environments. Therefore, we focus on the Resource-Limited Joint Multimodal Sentiment Reasoning and Classification task, JMSRC, which simultaneously performs multimodal sentiment reasoning chain generation and sentiment classification only with a lightweight model. We propose a Multimodal Chain-of-Thought Reasoning Distillation model, MulCoT-RD, designed for JMSRC that employs a "Teacher-Assistant-Student" distillation paradigm to address deployment constraints in resource-limited environments. We first leverage a high-performance Multimodal Large Language Model (MLLM) to generate the initial reasoning dataset and train a medium-sized assistant model with a multi-task learning mechanism. A lightweight student model is jointly trained to perform efficient multimodal sentiment reasoning generation and classification. Extensive experiments on four datasets demonstrate that MulCoT-RD with only 3B parameters achieves strong performance on JMSRC, while exhibiting robust generalization and enhanced interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。