用生成式知识提示提升多模态情感计算模型性能
Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting
- 提出结合生成知识提示与微调的混合增强策略
- 在多个数据集上显著提升多模态情感识别效果
- 为模型优化提供架构与数据特性分析依据
多模态情感计算(MAC)旨在通过融合文本、视频和音频等多源信息来识别和理解人类情绪。近年来,多模态大语言模型(MLLMs)为处理和对齐跨模态信息提供了统一框架,显著推动了该领域发展。然而,实际应用中仍存在任务复杂性导致的性能波动,以及对模型架构设计和数据特性影响缺乏深入理解等问题。为此,我们系统评估了当前主流开源的可同时处理音视频文本的MLLMs,在多个标准MAC数据集上的表现。评估不仅比较了模型性能,还通过分析模型结构与数据属性的影响,为优化提供可操作洞察。此外,我们提出一种新型混合策略,将生成式知识提示与监督微调相结合,以增强MLLMs的情感计算能力。实验结果表明,该方法在多种MAC任务中均取得显著提升,为未来研究与发展提供了可行路径。代码已开源:https://github.com/LuoMSen/MLLM-MAC。
原文摘要 · Abstract (English)
Multimodal Affective Computing (MAC) aims to recognize and interpret human emotions by integrating information from diverse modalities such as text, video, and audio. Recent advancements in Multimodal Large Language Models (MLLMs) have significantly reshaped the landscape of MAC by offering a unified framework for processing and aligning cross-modal information. However, practical challenges remain, including performance variability across complex MAC tasks and insufficient understanding of how architectural designs and data characteristics impact affective analysis. To address these gaps, we conduct a systematic benchmark evaluation of state-of-the-art open-source MLLMs capable of concurrently processing audio, visual, and textual modalities across multiple established MAC datasets. Our evaluation not only compares the performance of these MLLMs but also provides actionable insights into model optimization by analyzing the influence of model architectures and dataset properties. Furthermore, we propose a novel hybrid strategy that combines generative knowledge prompting with supervised fine-tuning to enhance MLLMs' affective computing capabilities. Experimental results demonstrate that this integrated approach significantly improves performance across various MAC tasks, offering a promising avenue for future research and development in this field. Our code is released on https://github.com/LuoMSen/MLLM-MAC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。