arXiv:2509.14738cs.CL2025-09EMNLP被引 1

构建统一视觉语言数据集,提升多模态理解与生成协同能力

UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets

  • 提出UnifiedVisual框架,整合理解与生成任务
  • 构建24万样本高质量数据集,支持跨模态推理
  • 适合研究多模态大模型与数据构造的学者使用

统一视觉大语言模型(VLLMs)在多模态理解与生成方面取得显著进展,推动了视觉问答与文本引导图像合成等应用的发展。然而,统一VLLM的进步受限于缺乏能充分发挥理解与生成协同潜力的数据集。现有数据集通常孤立处理理解与生成任务,制约了统一VLLM性能。为此,我们提出新型数据集构建框架UnifiedVisual,推出高质量数据集UnifiedVisual-240K,精心设计以促进多模态理解与生成的相互增强。该数据集无缝融合多样化的视觉与文本输入输出,支持全面的跨模态推理与精准的文本到图像对齐。涵盖广泛的任务与数据源,确保丰富多样性并弥补先前资源的关键缺陷。大量实验表明,基于UnifiedVisual-240K训练的模型在多种任务上均表现优异,尤其展现出理解与生成之间的显著相互强化,进一步验证了本框架与数据集的有效性。我们认为UnifiedVisual是推动统一VLLMs发展的新突破口,有望释放其全部潜力。代码与数据集已开源:https://github.com/fnlp-vision/UnifiedVisual。

原文摘要 · Abstract (English)

Unified vision large language models (VLLMs) have recently achieved impressive advancements in both multimodal understanding and generation, powering applications such as visual question answering and text-guided image synthesis. However, progress in unified VLLMs remains constrained by the lack of datasets that fully exploit the synergistic potential between these two core abilities. Existing datasets typically address understanding and generation in isolation, thereby limiting the performance of unified VLLMs. To bridge this critical gap, we introduce a novel dataset construction framework, UnifiedVisual, and present UnifiedVisual-240K, a high-quality dataset meticulously designed to facilitate mutual enhancement between multimodal understanding and generation. UnifiedVisual-240K seamlessly integrates diverse visual and textual inputs and outputs, enabling comprehensive cross-modal reasoning and precise text-to-image alignment. Our dataset encompasses a wide spectrum of tasks and data sources, ensuring rich diversity and addressing key shortcomings of prior resources. Extensive experiments demonstrate that models trained on UnifiedVisual-240K consistently achieve strong performance across a wide range of tasks. Notably, these models exhibit significant mutual reinforcement between multimodal understanding and generation, further validating the effectiveness of our framework and dataset. We believe UnifiedVisual represents a new growth point for advancing unified VLLMs and unlocking their full potential. Our code and datasets is available at https://github.com/fnlp-vision/UnifiedVisual.

多模态数据集构建视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。