调换图文顺序能影响大模型推理表现,关键在符合思维逻辑。
Image First or Text First? Optimising the Sequencing of Modalities in Large Language Model Prompting and Reasoning Tasks
- 按推理流程排列图文顺序,优于固定图文顺序。
- 简单任务中顺序差异导致准确率显著变化。
- 适合教育、医疗等需多步推理的跨模态应用。
本文研究多模态提示中图像与文本的呈现顺序如何影响大语言模型(LLMs)的推理表现。通过在三个商用LLM上进行实证评估发现,模态顺序对性能有显著影响,尤其在不同复杂度的任务中。对于仅涉及单张图像的简单任务,顺序变化明显影响准确率;而在包含多张图像和复杂推理步骤的复杂任务中,顺序影响减弱,可能源于任务本身认知负荷增加。研究还表明,问题/提示结构至关重要:在嵌套式与多步推理任务中,模态顺序对模型表现起关键作用。尽管LLM在推理初期表现良好,但难以重新整合早期信息,凸显其在多跳推理中的局限性。因此,将模态顺序与推理逻辑流程对齐,比单纯关注模态顺序更为重要。这些发现为优化多模态提示设计提供了实用指导,可广泛应用于教育、医学影像及跨模态学习等领域。
原文摘要 · Abstract (English)
This paper examines how the sequencing of images and text within multi-modal prompts influences the reasoning performance of large language models (LLMs). We performed empirical evaluations using three commercial LLMs. Our results demonstrate that the order in which modalities are presented can significantly affect performance, particularly in tasks of varying complexity. For simpler tasks involving a single image, modality sequencing had a clear impact on accuracy. However, in more complex tasks involving multiple images and intricate reasoning steps, the effect of sequencing diminished, likely due to the increased cognitive demands of the task. Our findings also highlight the importance of question/prompt structure. In nested and multi-step reasoning tasks, modality sequencing played a key role in shaping model performance. While LLMs excelled in the initial stages of reasoning, they struggled to re-incorporate earlier information, underscoring the challenges of multi-hop reasoning within transformer architectures. This suggests that aligning the sequence of modalities with the logical flow of reasoning steps is more critical than modality order alone. These insights offer valuable implications for improving multi-modal prompt design, with broader applications across fields such as education, medical imaging, and cross-modal learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。