arXiv:2507.16746cs.CVcs.CL2025-07被引 79

构建大规模图文交错推理数据集,提升模型复杂问题求解能力

Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning

  • 设计包含18万+样本的图文交织推理数据集
  • 微调后模型在测试集准确率提升12%,部分基准提高13%
  • 适用于需要视觉推理的科研、机器人与逻辑游戏场景

人类在解决复杂问题时常借助图示或草图。训练多模态模型实现类似能力(即视觉链式思维,Visual CoT)面临两大挑战:一是现有视觉CoT性能不佳,制约强化学习;二是缺乏高质量训练数据。本文提出Zebra-CoT,一个包含182,384个样本的大规模多样化数据集,涵盖逻辑连贯的图文交错推理过程。聚焦四类适合绘图推理的任务:几何、物理与算法等科学问题;视觉搜索、拼图等2D任务;3D多跳推理、具身与机器人规划等3D任务;以及棋类等策略逻辑题。在Zebra-CoT上微调Anole-7B模型,测试集准确率提升12%,标准VLM基准测评最高增益达13%。微调Bagel-7B则生成高质量图文交替推理链,验证数据集对多模态推理能力的促进作用。数据集与模型已开源,支持视觉CoT研究与评估。

原文摘要 · Abstract (English)

Humans often use visual aids, for example diagrams or sketches, when solving complex problems. Training multimodal models to do the same, known as Visual Chain of Thought (Visual CoT), is challenging due to: (1) poor off-the-shelf visual CoT performance, which hinders reinforcement learning, and (2) the lack of high-quality visual CoT training data. We introduce $\textbf{Zebra-CoT}$, a diverse large-scale dataset with 182,384 samples, containing logically coherent interleaved text-image reasoning traces. We focus on four categories of tasks where sketching or visual reasoning is especially natural, spanning scientific questions such as geometry, physics, and algorithms; 2D visual reasoning tasks like visual search and jigsaw puzzles; 3D reasoning tasks including 3D multi-hop inference, embodied and robot planning; visual logic problems and strategic games like chess. Fine-tuning the Anole-7B model on the Zebra-CoT training corpus results in an improvement of +12% in our test-set accuracy and yields up to +13% performance gain on standard VLM benchmark evaluations. Fine-tuning Bagel-7B yields a model that generates high-quality interleaved visual reasoning chains, underscoring Zebra-CoT's effectiveness for developing multimodal reasoning abilities. We open-source our dataset and models to support development and evaluation of visual CoT.

视觉推理多模态链式思维数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。