测试大模型在多模态情感分析中的表现,发现效果不佳且速度慢。
Exploring Large Language Models for Multimodal Sentiment Analysis: Challenges, Benchmarks, and Future Directions
- 构建新基准评估大模型在图文情感分析中的能力
- 大模型准确率低、推理慢,不如传统方法
- 适合关注多模态大模型局限的研究者
多模态方面情感分析(MABSA)旨在从文本和图像等多模态信息中提取方面词及其情感极性。尽管传统监督学习方法在此任务上表现有效,但大语言模型(LLMs)在该任务上的适应性仍不明确。近年来,Llama2、LLaVA和ChatGPT等大模型在通用任务中展现强大能力,但在复杂精细场景如MABSA中的表现尚待探索。本研究对大模型在MABSA中的适用性进行了全面考察,构建了基准测试以评估其性能,并与当前最先进的监督学习方法进行对比。实验表明,尽管大模型具备多模态理解潜力,但在准确性与推理时间方面面临显著挑战,难以达到满意结果。基于此,我们讨论了现有大模型的局限性,并提出了未来提升多模态情感分析能力的研究方向。
原文摘要 · Abstract (English)
Multimodal Aspect-Based Sentiment Analysis (MABSA) aims to extract aspect terms and their corresponding sentiment polarities from multimodal information, including text and images. While traditional supervised learning methods have shown effectiveness in this task, the adaptability of large language models (LLMs) to MABSA remains uncertain. Recent advances in LLMs, such as Llama2, LLaVA, and ChatGPT, demonstrate strong capabilities in general tasks, yet their performance in complex and fine-grained scenarios like MABSA is underexplored. In this study, we conduct a comprehensive investigation into the suitability of LLMs for MABSA. To this end, we construct a benchmark to evaluate the performance of LLMs on MABSA tasks and compare them with state-of-the-art supervised learning methods. Our experiments reveal that, while LLMs demonstrate potential in multimodal understanding, they face significant challenges in achieving satisfactory results for MABSA, particularly in terms of accuracy and inference time. Based on these findings, we discuss the limitations of current LLMs and outline directions for future research to enhance their capabilities in multimodal sentiment analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。