用军事指挥策略提升多模态模型的讽刺识别能力
Commander-GPT: Fully Unleashing the Sarcasm Detection Capability of Multi-Modal Large Language Models
- 将讽刺识别拆解为六项子任务,由中心模型分配最优大模型处理
- 在MMSD数据集上F1分数提升19.3%,无需微调或推理说明
- 适合需要高精度多模态讽刺检测的研究与应用
讽刺识别是自然语言处理中的关键方向,传统单模态方法因讽刺的隐晦性难以取得理想效果。近年来研究转向多模态方法,但如何有效利用多模态信息仍具挑战。本文提出一种名为Commander-GPT的创新框架,借鉴军事指挥思想,将讽刺识别任务分解为六个子任务,由中央决策者分配最合适的大型语言模型分别处理,最终聚合结果判断讽刺。在MMSD和MMSD 2.0数据集上,结合四种多模态大模型与六种提示策略进行实验,结果表明该方法达到当前最优性能,F1分数提升19.3%,且无需微调或真实推理理由。
原文摘要 · Abstract (English)
Sarcasm detection, as a crucial research direction in the field of Natural Language Processing (NLP), has attracted widespread attention. Traditional sarcasm detection tasks have typically focused on single-modal approaches (e.g., text), but due to the implicit and subtle nature of sarcasm, such methods often fail to yield satisfactory results. In recent years, researchers have shifted the focus of sarcasm detection to multi-modal approaches. However, effectively leveraging multi-modal information to accurately identify sarcastic content remains a challenge that warrants further exploration. Leveraging the powerful integrated processing capabilities of Multi-Modal Large Language Models (MLLMs) for various information sources, we propose an innovative multi-modal Commander-GPT framework. Inspired by military strategy, we first decompose the sarcasm detection task into six distinct sub-tasks. A central commander (decision-maker) then assigns the best-suited large language model to address each specific sub-task. Ultimately, the detection results from each model are aggregated to identify sarcasm. We conducted extensive experiments on MMSD and MMSD 2.0, utilizing four multi-modal large language models and six prompting strategies. Our experiments demonstrate that our approach achieves state-of-the-art performance, with a 19.3% improvement in F1 score, without necessitating fine-tuning or ground-truth rationales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。