arXiv:2504.16427cs.CLcs.AI2025-04NeurIPS被引 7

构建首个多模态语言理解综合评测基准,揭示大模型在语义理解上的局限。

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark

  • 设计覆盖6类语义维度的多模态数据集,涵盖真实与模拟场景
  • 主流模型在多模态语义任务上准确率仅60%~70%,仍存显著差距
  • 适合研究多模态理解、人机交互及大模型认知能力的学者使用

多模态语言分析通过融合多种模态提升对人类对话深层语义的理解。尽管该领域发展迅速,但对多模态大语言模型(MLLMs)在认知级语义理解方面的能力研究仍不足。本文提出MMLA,一个专为填补此空白而设计的综合性评测基准。MMLA包含超过6.1万条来自真实与设定场景的多模态话语,覆盖意图、情感、对话行为、情感倾向、表达风格和沟通行为六大核心维度。我们采用零样本推理、监督微调和指令微调三种方法,评估八种主流的LLMs与MLLMs。大量实验表明,即使经过微调,模型准确率也仅为60%~70%,凸显当前多模态大模型在理解复杂人类语言方面的显著局限。我们认为MMLA将成为探索大模型在多模态语言分析中潜力的坚实基础,并为该领域提供宝贵资源。数据集与代码已开源至https://github.com/thuiar/MMLA。

原文摘要 · Abstract (English)

Multimodal language analysis is a rapidly evolving field that leverages multiple modalities to enhance the understanding of high-level semantics underlying human conversational utterances. Despite its significance, little research has investigated the capability of multimodal large language models (MLLMs) to comprehend cognitive-level semantics. In this paper, we introduce MMLA, a comprehensive benchmark specifically designed to address this gap. MMLA comprises over 61K multimodal utterances drawn from both staged and real-world scenarios, covering six core dimensions of multimodal semantics: intent, emotion, dialogue act, sentiment, speaking style, and communication behavior. We evaluate eight mainstream branches of LLMs and MLLMs using three methods: zero-shot inference, supervised fine-tuning, and instruction tuning. Extensive experiments reveal that even fine-tuned models achieve only about 60%~70% accuracy, underscoring the limitations of current MLLMs in understanding complex human language. We believe that MMLA will serve as a solid foundation for exploring the potential of large language models in multimodal language analysis and provide valuable resources to advance this field. The datasets and code are open-sourced at https://github.com/thuiar/MMLA.

多模态语言理解大模型评测认知语义

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。