构建百万级跨学科视觉生成数据集,提升知识密集型图表生成准确性
DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing

- 设计多阶段流水线合成结构化学术图像与编辑指令
- 在10个学科领域覆盖120万样本,显著提升知识图生成精度
- 适合研究多模态生成、教育科技与可验证视觉创作的学者
当前图像生成与编辑模型虽能产出视觉美观的自然图像,但在生成依赖学科知识、符号结构与精确空间关系的知识密集型图表时仍不可靠。我们提出DisciplineGen-1M,一个百万级跨学科视觉生成与编辑数据集,涵盖数学、物理、化学、生物、地理、计算机科学、经济学、历史、音乐和体育共10个领域,总计120万样本。通过结构化渲染、OCR编辑、专用程序合成与大规模文本到图像过滤等可扩展流程,生成带标注的图像、描述文本、编辑指令及具有可控语义差异的配对图像。基于该数据集,我们进一步构建了融合学科知识的推理-生成模型,用于图文生成与图像编辑。在GenExam与GRADE等学科基准测试中,性能显著优于开源基线;在WISE与RISE等通用推理基准上也展现良好泛化能力。结果表明,大规模结构化学术视觉数据是实现从审美合理性迈向可验证知识驱动视觉创作的关键。我们将公开发布数据集、模型及数据构建源码,以保障可复现性并推动后续研究。
原文摘要 · Abstract (English)
Recent image generation and editing models can produce visually appealing natural images, yet they remain unreliable when the target image is a knowledge-intensive diagram whose correctness depends on disciplinary concepts, symbolic structure, and precise spatial relations. We introduce DisciplineGen-1M, a million-scale multidisciplinary dataset that supports text-to-image generation and image editing. It contains 1.2M samples spanning mathematics, physics, chemistry, biology, geography, computer science, economics, history, music, and sports. To construct the dataset, we design a scalable framework that combines structured rendering, OCR-based editing, specialized programmatic synthesis, and large-scale text-to-image filtering. These pipelines produce captions, editing instructions, structured annotations, and paired images with controllable semantic differences. Building on DisciplineGen-1M, we further introduce a discipline-informed reasoning-generation model for both text-to-image generation and image editing. Experiments on discipline-related benchmarks, GenExam and GRADE, show substantial improvements over open-source baselines, while evaluations on general reasoning-informed benchmarks, WISE and RISE, further indicate broader transfer. The results suggest that large-scale structured academic visual data is a key ingredient for moving image generation from aesthetic plausibility toward verifiable knowledge-grounded visual creation. We will publicly release our dataset, model, and source code of the data curation pipeline to ensure reproducibility and benefit future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。