arXiv:2505.16193cs.CLcs.CV2025-05ICML被引 6

优化提示示例配置,显著提升多模态模型的情感感知能力

An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception Capability

  • 通过优化示例的检索、呈现和分布三要素,增强模型情感理解
  • 在六个数据集上相较零样本基准平均提升15.9%准确率
  • 发现并纠正了模型固有的情感预测偏差,适合相关应用开发者

多模态大模型(MLLMs)的发展使各类多模态任务可在零样本范式下完成,避免了模型微调成本,成为实际应用主流。然而,多模态情感分析(MSA)这一通用人工智能的关键挑战仍难以适用该范式,其零样本表现不佳,引发对MLLM情感感知能力的质疑。本文将零样本范式扩展至上下文学习(ICL),深入研究示例配置策略,验证了MLLM确实具备情感感知能力。具体考察了示例的检索、呈现与分布三个关键因素,并进行了系统优化。同时发现并有效缓解了模型固有的情感预测偏差。三项策略协同作用,在六组MSA数据集上相较零样本范式平均提升15.9%准确率,相较随机ICL基线提升11.2%。

原文摘要 · Abstract (English)

The advancements in Multimodal Large Language Models (MLLMs) have enabled various multimodal tasks to be addressed under a zero-shot paradigm. This paradigm sidesteps the cost of model fine-tuning, emerging as a dominant trend in practical application. Nevertheless, Multimodal Sentiment Analysis (MSA), a pivotal challenge in the quest for general artificial intelligence, fails to accommodate this convenience. The zero-shot paradigm exhibits undesirable performance on MSA, casting doubt on whether MLLMs can perceive sentiments as competent as supervised models. By extending the zero-shot paradigm to In-Context Learning (ICL) and conducting an in-depth study on configuring demonstrations, we validate that MLLMs indeed possess such capability. Specifically, three key factors that cover demonstrations' retrieval, presentation, and distribution are comprehensively investigated and optimized. A sentimental predictive bias inherent in MLLMs is also discovered and later effectively counteracted. By complementing each other, the devised strategies for three factors result in average accuracy improvements of 15.9% on six MSA datasets against the zero-shot paradigm and 11.2% against the random ICL baseline.

多模态情感分析上下文学习模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。