arXiv:2510.21828cs.CVcs.CL2025-10ACL被引 1

构建多模态关系知识数据集,提升小模型在抽象推理上的表现。

Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images

  • 自动生成含结构化关系的多模态图像数据,支持思维链训练。
  • 3B/7B小模型经训练后超越GPT-4o在抽象推理任务中的表现。
  • 适用于需要复杂多模态推理的研究者与开发者。

当前多模态大模型在理解视觉模态的抽象信息方面仍面临挑战。其中,以节点-边形式表示多模态实体间抽象关系的多模态关系知识(MMRK)尚未得到充分研究。本文提出一种自动化的STAR数据生成引擎,用于合成包含MMRK的多模态指令数据,支持各类抽象推理任务的思维链训练;并设计了一个两阶段能力增强训练框架,配套评估协议。基于此,我们构建了包含64K高质量样本的STAR-64K数据集,在5个开源多模态大模型上进行实验。结果表明,该框架使3B/7B小模型在抽象推理任务中显著优于GPT-4o。同时,我们对设计有效性、数据可迁移性与可扩展性进行了深入分析。

原文摘要 · Abstract (English)

Understanding and reasoning with abstractive information from the visual modality presents significant challenges for current multi-modal large language models (MLLMs). Among the various forms of abstractive information, Multi-Modal Relational Knowledge (MMRK), which represents abstract relational structures between multi-modal entities using node-edge formats, remains largely under-explored. In particular, STructured and Abstractive Reasoning (STAR) on such data has received little attention from the research community. To bridge the dual gaps in large-scale high-quality data and capability enhancement methodologies, this paper makes the following key contributions: (i). An automatic STAR data engine capable of synthesizing images with MMRK to build multi-modal instruction data with reliable chain-of-thought thinking for various STAR tasks and (ii). A comprehsive two-stage capability enhancement training framework, accompanied by a suite of evaluation protocols tailored to different STAR tasks. Based upon these contributions, we introduce STAR-64K, a dataset comprising 64K high-quality multi-modal instruction samples, and conduct experiments across 5 open-source MLLMs. Experimental results show that our two-stage enhancement framework enables smaller 3B/7B models to significantly outperform GPT-4o in STAR. Additionally, we provide in-depth analysis regarding the effectiveness of various designs, data transferability, and scalability.

多模态抽象推理知识图谱小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。