首个支持多轮跨模态理解与生成的基准数据集,推动视觉记忆与上下文推理研究。
WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation
- 构建包含100K交互样本的大规模数据集,支持多轮图文对话与跨模态任务。
- 在WEAVEBench上测试显示,模型在多轮生成与视觉记忆方面仍有明显短板。
- 适用于研究视觉记忆、多轮交互和上下文感知图像编辑的团队与开发者。
统一多模态模型(UMMs)在视觉理解与生成方面取得了显著进展,但现有数据集和基准主要聚焦单轮交互,未能捕捉真实世界图像创作与编辑中多轮、上下文依赖的特性。为此,我们提出WEAVE,首个面向上下文内交错跨模态理解与生成的完整套件。WEAVE-100k是一个大规模数据集,包含100K个交错样本,覆盖超过37万次对话轮次和50万张图像,涵盖需基于历史上下文推理的理解、编辑与生成任务。WEAVEBench是基于480张图像构建的人工标注基准,包含100个任务,采用结合参考图与原始图+编辑指令的混合视觉语言模型(VLM)评判框架,评估模型在多轮生成、视觉记忆与世界知识推理方面的表现。实验表明,基于WEAVE-100k训练可使UMMs具备视觉理解、图像编辑及理解-生成协同能力,并催生涌现的视觉记忆能力;然而,在WEAVEBench上的广泛评估揭示了当前方法在多轮、上下文感知图像生成与编辑中的持续局限与挑战。我们相信WEAVE为多模态社区研究上下文内交错理解与生成提供了新视角与基础。
原文摘要 · Abstract (English)
Recent advances in unified multimodal models (UMMs) have enabled impressive progress in visual comprehension and generation. However, existing datasets and benchmarks focus primarily on single-turn interactions, failing to capture the multi-turn, context-dependent nature of real-world image creation and editing. To address this gap, we present WEAVE, the first suite for in-context interleaved cross-modality comprehension and generation. Our suite consists of two complementary parts. WEAVE-100k is a large-scale dataset of 100K interleaved samples spanning over 370K dialogue turns and 500K images, covering comprehension, editing, and generation tasks that require reasoning over historical context. WEAVEBench is a human-annotated benchmark with 100 tasks based on 480 images, featuring a hybrid VLM judger evaluation framework based on both the reference image and the combination of the original image with editing instructions that assesses models' abilities in multi-turn generation, visual memory, and world-knowledge reasoning across diverse domains. Experiments demonstrate that training on WEAVE-100k enables vision comprehension, image editing, and comprehension-generation collaboration capabilities. Furthermore, it facilitates UMMs to develop emergent visual-memory capabilities, while extensive evaluations on WEAVEBench expose the persistent limitations and challenges of current approaches in multi-turn, context-aware image generation and editing. We believe WEAVE provides a view and foundation for studying in-context interleaved comprehension and generation for multi-modal community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。