用空间知识图谱指导生成更符合常识的多模态数据。
Spatial Knowledge Graph-Guided Multimodal Synthesis
- 基于空间知识图谱自动构建空间关系,引导多模态合成。
- 生成的数据显著提升MLLM的空间感知与推理能力。
- 适合研究空间智能、多模态生成的开发者使用。
多模态大语言模型(MLLMs)虽有显著进步,但空间感知能力仍受限。为此,本文提出SKG2DATA,一种基于空间知识图谱(SKG)的多模态数据合成方法,实现从知识到数据的生成。该方法通过自动化流程构建捕捉方向与距离关系的结构化空间知识图谱,为扩散模型生成空间一致图像、MLLM生成对应文本提供精准引导。该框架支持可扩展、多样且真实的场景配置生成,克服人工标注的局限。实验表明,基于方向与距离等空间知识合成的数据,能显著增强MLLM的空间感知与推理能力,尽管对通用能力略有影响。代码已开源。
原文摘要 · Abstract (English)
Recent advances in Multimodal Large Language Models (MLLMs) have significantly enhanced their capabilities; however, their spatial perception abilities remain a notable limitation. To address this challenge, multimodal data synthesis offers a promising solution. Yet, ensuring that synthesized data adhere to spatial common sense is a non-trivial task. Our approach addresses this critical gap by providing a systematic framework for generating spatially coherent data. In this work, we introduce SKG2DATA, a novel multimodal synthesis approach guided by spatial knowledge graphs, grounded in the concept of knowledge-to-data generation. SKG2DATA employs an automated pipeline for constructing Spatial Knowledge Graph (SKG) that effectively captures human-like spatial cognition, including directional and distance relationships. These structured representations then serve as precise guidance for our integrated synthesis pipeline, where a diffusion model generates spatially-consistent images while a MLLM produces corresponding textual descriptions. The automated construction of SKG enables scalable generation of diverse yet realistic spatial configurations, overcoming the limitations of manual data collection and annotation. Extensive experiments demonstrate that data synthesized from diverse types of spatial knowledge, including direction and distance, enhance the spatial perception and reasoning abilities of MLLMs markedly, albeit with a slight cost to their general capabilities. We hope that the idea of knowledge-based data synthesis can advance the development of spatial intelligence. Code is available at https://github.com/zjunlp/Knowledge2Data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。