构建图结构自动数据生成框架,提升多模态跨模态推理能力
CRIT: Graph-Based Automatic Data Synthesis to Enhance Cross-Modal Multi-Hop Reasoning
- 基于图结构自动合成复杂跨模态推理任务数据
- 新数据集上顶尖模型仍表现不佳,验证任务难度
- 显著提升模型在SPIQA等基准上的多跳推理能力
现实世界推理常需跨模态整合信息,将文本上下文与视觉线索通过多跳过程关联。然而现有多模态评测集多依赖单张或一组图像,答案可仅从单一模态推断,无法体现真正的跨模态多跳推理。训练数据也缺乏交错图文内容,导致视觉语言模型(VLMs)频繁幻觉,推理过程缺乏视觉证据支撑。为此,我们提出CRIT,一个基于图的自动化数据合成管道,构建复杂跨模态推理任务的新数据集与评测基准。CRIT涵盖自然图像、视频及文本丰富来源,包含人工验证的测试集以确保评估可靠性。实验表明,即使最先进的模型在该基准上仍表现有限;而在CRIT上训练的模型在跨模态多跳推理上取得显著提升,尤其在SPIQA等标准多模态基准上表现增强。
原文摘要 · Abstract (English)
Real-world reasoning often requires combining information across modalities, connecting textual context with visual cues in a multi-hop process. Yet, most multimodal benchmarks fail to capture this ability: they typically rely on single images or set of images, where answers can be inferred from a single modality alone. This limitation is mirrored in the training data, where interleaved image-text content rarely enforces complementary, multi-hop reasoning. As a result, Vision-Language Models (VLMs) frequently hallucinate and produce reasoning traces poorly grounded in visual evidence. To address this gap, we introduce CRIT, a new dataset and benchmark built with a graph-based automatic pipeline for generating complex cross-modal reasoning tasks. CRIT consists of diverse domains ranging from natural images, videos, and text-rich sources, and includes a manually verified test set for reliable evaluation. Experiments on this benchmark reveal that even state-of-the-art models struggle on such reasoning tasks. Models trained on CRIT show significant gains in cross-modal multi-hop reasoning, including strong improvements on SPIQA and other standard multimodal benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。