用任务提示提升视觉模型数学推理能力
Beyond Captioning: Task-Specific Prompting for Improved VLM Performance in Mathematical Reasoning
- 针对数学任务设计特定提示,替代通用描述生成
- 大模型在数学问题上直接生成描述效果差,随机性高
- 任务提示显著优于传统图文描述,适合数学推理场景
视觉语言模型(VLMs)在需要视觉与推理结合的任务中表现突出,如图像检索和视觉问答(VQA)。然而,在几何推理、代数求解和计数等任务中仍存在显著瓶颈,主要源于多模态信息融合困难及对几何类任务理解不准。现有研究认为在VQA前加入图文描述生成可提升性能,但本文发现该方法不具备泛化性,尤其在以下游问答任务为主训练的大规模模型上,数学相关任务表现接近随机。为此,我们提出任务导向的提示策略,通过在提示中注入具体任务引导信息,有效提升了模型在几何、代数和计数任务上的表现。实验表明,该方法优于直接使用图文描述的基准方案,为数学密集型视觉任务提供了更优解决方案。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have transformed tasks requiring visual and reasoning abilities, such as image retrieval and Visual Question Answering (VQA). Despite their success, VLMs face significant challenges with tasks involving geometric reasoning, algebraic problem-solving, and counting. These limitations stem from difficulties effectively integrating multiple modalities and accurately interpreting geometry-related tasks. Various works claim that introducing a captioning pipeline before VQA tasks enhances performance. We incorporated this pipeline for tasks involving geometry, algebra, and counting. We found that captioning results are not generalizable, specifically with larger VLMs primarily trained on downstream QnA tasks showing random performance on math-related challenges. However, we present a promising alternative: task-based prompting, enriching the prompt with task-specific guidance. This approach shows promise and proves more effective than direct captioning methods for math-heavy problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。