构建大规模视觉抽象链数据集,让模型学会图文混合推理。
CoVA-SFT: A Large-Scale Dataset for Chain of Visual Abstractions
- 设计51.9万步多模态推理数据,含5类布局与17项复杂任务
- 微调模型在基准测试中平均性能超基线2倍以上
- 适合研究多模态推理、视觉思维链的学者与工程师
思维链(CoT)推理显著提升了大语言模型的能力,使其能将问题分解为中间步骤。尽管在语言任务中效果显著,纯文本思维链难以有效处理视觉问题。现有架构虽可处理视觉输入,但缺乏大规模、多步骤、自修正的训练数据来教导模型如何构建和维护解决纯文本推理问题时的内部视觉工作空间。为此,我们提出CoVA-SFT,一个包含51.9K样本、超过222K个多模态推理步骤的结构化语料库,涵盖5种不同布局类型和17个复杂任务,并配套推出CoVA-Bench基准测试集,包含1,700个独立测试样本,支持可复现评估。通过提供明确的推理表述、代理式渲染和验证循环,CoVA-SFT促使多模态语言模型在文本与视觉抽象间交替生成。实验证明,基于CoVA-SFT微调的模型在CoVA-Bench上平均性能超越所有交错式思维链基线两倍以上,但仍落后于强文本型思维链基线,揭示了未来研究的挑战。
原文摘要 · Abstract (English)
Chain-of-thought (CoT) reasoning has dramatically improved large language models (LLMs) by allowing them to decompose problems into intermediate steps. While CoT is widely effective for linguistic tasks, text-only CoT forces models to serialize visual problems into awkward prose. Although architectural solutions exist to process visual inputs, the community lacks a massive, multi-step, self-corrected dataset to teach models how to build and maintain internal visual workspaces when solving purely textual reasoning problems. To address this limitation, we introduce CoVA-SFT, a highly structured corpus of 51.9K samples containing over 222K multimodal reasoning steps across 5 distinct layout families and 17 complex tasks, and CoVA-Bench, a companion benchmark of 1,700 held-out test samples spanning the same tasks for reproducible evaluation. By providing explicit rationale formulations, agentic renderings, and verification loops, CoVA-SFT teaches multimodal language models to interleave text and visual abstractions. We validate the dataset by demonstrating that models fine-tuned on CoVA-SFT outperform all interleaved CoT baselines by more than 2x on average on CoVA-Bench, though they still fall short of strong text-only CoT baselines, highlighting open challenges for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。