文本主导的多模态意图识别数据集存在严重偏见,影响模型评估公平性。
Text Takes Over: A Study of Modality Bias in Multimodal Intent Detection
- 发现文本仅模型在多模态意图识别中表现优于多数融合模型。
- 超过90%样本依赖文本信息,导致多模态模型难以发挥优势。
- 提出去偏框架,移除超半数样本后模型性能大幅下降。
多模态数据(文本、音频、视觉)的兴起为意图检测等任务带来新机遇。本研究评估了大语言模型(LLM)与非LLM(包括纯文本和多模态模型)在多模态意图检测中的表现。结果显示,纯文本模型Mistral-7B在MIntRec-1和MIntRec2.0数据集上分别比多数多模态模型高出约9%和4%。这一优势源于数据集中超过90%的样本需依赖文本输入(单独或与其他模态结合)才能正确分类。人工评估进一步验证了这种模态偏见。为此,我们提出一个去偏框架,去偏后,MIntRec-1中超过70%、MIntRec2.0中超过50%的样本被移除,导致所有模型性能显著下降,其中小型多模态融合模型准确率下降超过50%-60%。通过实证分析不同模态的上下文相关性,揭示了多模态数据集中的模态偏见问题,强调构建无偏数据集对有效评估多模态模型的重要性。
原文摘要 · Abstract (English)
The rise of multimodal data, integrating text, audio, and visuals, has created new opportunities for studying multimodal tasks such as intent detection. This work investigates the effectiveness of Large Language Models (LLMs) and non-LLMs, including text-only and multi-modal models, in the multimodal intent detection task. Our study reveals that Mistral-7B, a text-only LLM, outperforms most competitive multimodal models by approximately 9% on MIntRec-1 and 4% on MIntRec2.0 datasets. This performance advantage comes from a strong textual bias in these datasets, where over 90% of the samples require textual input, either alone or in combination with other modalities, for correct classification. We confirm the modality bias of these datasets via human evaluation, too. Next, we propose a framework to debias the datasets, and upon debiasing, more than 70% of the samples in MIntRec-1 and more than 50% in MIntRec2.0 get removed, resulting in significant performance degradation across all models, with smaller multimodal fusion models being the most affected with an accuracy drop of over 50 - 60%. Further, we analyze the context-specific relevance of different modalities through empirical analysis. Our findings highlight the challenges posed by modality bias in multimodal intent datasets and emphasize the need for unbiased datasets to evaluate multimodal models effectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。