构建高质量图文交错生成数据集与评估框架,助力大模型提升交互式多模态能力。
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
- 构建180万样本的InterSyn数据集,通过自评迭代优化确保质量
- 在20万样本下性能显著提升,100万以上样本进一步改善图文协同效果
- 提出可解释的SynJudge评估体系,适合资源有限的研究者使用
大型多模态模型在多模态理解与生成方面取得进展,但在生成紧密交错的图文输出方面仍存在不足,主要源于现有训练数据集规模小、质量低且指令多样性不足。为此,我们提出InterSyn数据集,包含180万个多模态样本,具备三大特性:(1)大规模;(2)高质量,基于自评估与迭代精炼(SEIR)方法实现严格自动化质量控制;(3)丰富的指令多样性,基于人类偏好设计多样化问题模板,覆盖3500个主题层级。该数据集特别适用于训练具备交互式图文生成能力的LMM。为评估模型表现,我们提出SynJudge自动评估系统,其评分与人工评判高度一致,输出四项可解释指标:文本内容完整性(TCC)、图像内容完整性(ICC)、图像质量(IQ)和图文协同性(ITS)。实验结果表明,在最大达20万样本的子集上,2.5万至5万样本即带来显著提升,扩展至10万和20万样本后,TCC、ICC和尤其ITS持续增长,验证了InterSyn的可扩展性与高效性——即使小规模数据也能实现显著改进,适合不同算力条件的研究者使用。
原文摘要 · Abstract (English)
Recent advancements in Large Multimodal Models (LMMs) have significantly improved multimodal understanding and generation. However, these models still struggle to generate tightly interleaved image-text outputs, primarily due to the limited scale, quality, and instructional richness of current training datasets. To address this, we introduce InterSyn, a dataset that features: (1) large scale, comprising 1.8M multimodal samples; (2) high quality, supported by our proposed Self-Evaluation with Iterative Refinement (SEIR) method for rigorous automated quality refinement; (3) rich instructional diversity, ensured through diverse well-designed question templates, based on human preferences and covering a 3500-topic hierarchy. These characteristics make InterSyn particularly well-suited for training LMMs in interactive image-text generation capabilities. To evaluate the capabilities, we propose SynJudge, a reliable automatic evaluator that aligns closely with human judge and outputs four interpretable scores: Text Content Completeness (TCC), Image Content Completeness (ICC), Image Quality (IQ), and Image-Text Synergy (ITS). These scores are complementary, covering both content and quality as well as cross-modal interaction, thereby forming a comprehensive evaluation framework. Experimental results on InterSyn subsets of up to 200K samples show that 25K-50K already yield substantial improvements, while scaling to 100K/200K brings further gains in TCC, ICC, and especially ITS, highlighting InterSyn's: (1) scalability, as performance consistently improves with more data; (2) efficiency, as significant gains are achievable even with smaller subsets, making it accessible to researchers with varying computational resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。