让大模型先描述多模态输入,再回答问题,提升理解与推理能力。
Chain-of-Description: What I can understand, I can put into words
- 先详细描述输入内容,再作答,增强模型理解。
- 音频任务提升近4%,视觉难题部分提升5.3%。
- 适合需要精准多模态理解的科研与应用开发人员。
本文提出一种针对多模态大模型的新策略——链式描述(Chain-of-Description, CoD)提示方法。该方法要求模型在回答问题前,先对多模态输入进行详细描述。在Qwen2-Audio、Qwen2-VL和Qwen2.5-VL等模型上测试表明,相比标准提示方法,CoD Prompting显著提升性能:在音频基准AIR-Bench-Chat的语音类别中提升近4%,在视觉基准MMMU_Pro的高难度部分提升5.3%。消融实验进一步验证了该方法的有效性。
原文摘要 · Abstract (English)
In this paper, we propose a novel strategy defined as Chain-of-Description (CoD) Prompting, tailored for Multi-Modal Large Language Models. This approach involves having the model first provide a detailed description of the multi-modal input before generating an answer to the question. When applied to models such as Qwen2-Audio, Qwen2-VL, and Qwen2.5-VL, CoD Prompting significantly enhances performance compared to standard prompting methods. This is demonstrated by nearly a 4\% improvement in the speech category of the audio benchmark AIR-Bench-Chat and a 5.3\% improvement in the hard-level portion of the vision benchmark MMMU\_Pro. Our ablation study further validates the effectiveness of CoD Prompting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。