arXiv:2606.30124cs.CV2026-06中稿 · ECCV

构建科学图像生成的大型数据集与评测基准,提升模型科学推理能力。

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation

论文配图:SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation
图 1 · 摘自论文原文
  • 基于符号学三元组构建科学图像生成的三维度框架:实体结构、科学过程、科学规律。
  • 推出包含8.2万张图像-文本对的SciIR-82k数据集,引入科学推理链增强逻辑性。
  • 提出可细粒度验证的SciIR-Bench评测体系,适合研究科学图像生成与逻辑推理的学者。

尽管文本到图像(T2I)模型在生成逼真视觉内容方面取得显著进展,但在科学图像所需的严格语义对齐与逻辑推理方面仍存在不足。受皮尔斯符号学三元组启发,我们提出科学图像推理(SciIR),为科学图像生成的训练与评估提供综合性资源。将科学推理形式化为三个核心维度:实体结构(图标)、科学过程(指示符)、科学定律(符号)。为解决科学图像生成训练数据稀缺问题,我们精心构建了包含超过80,000个高质量图像-文本对的SciIR-82k数据集,其按符号学维度分层组织,并引入科学推理链(Sci-RCoT)以显式建模底层视觉逻辑。针对评估,我们设计了与三重符号维度一致的SciIR-Bench,采用原子检查表将结果导向的科学准确性转化为过程导向、可验证的细粒度问题。大量实验揭示当前模型在科学推理能力上的显著缺陷。通过在SciIR-82k上微调,我们开发出Qwen-Image-SciIR模型,在SciIR-Bench上得分从35%提升至43%,为未来科学图像生成的发展奠定坚实基础。

原文摘要 · Abstract (English)

While Text-to-Image (T2I) models have shown remarkable success in generating photorealistic visual content, they still struggle with the rigorous semantic alignment and logical reasoning required for scientific imagery. Inspired by Peirce's Semiotic Triad, we introduce Scientific Image Reasoning (SciIR), a comprehensive resource for training and evaluation of scientific image generation. We formalize scientific reasoning into three core dimensions: Entity Structure (Icon), Scientific Process (Index), and Scientific Law (Symbol). Specifically, to overcome the scarcity of training data in scientific image generation, we elaborately create SciIR-82k, a large-scale dataset containing over 80,000 high-quality scientific image-text pairs from cutting-edge publications. The dataset is hierarchically organized according to the semiotic dimensions and incorporates a Scientific Reasoning Chain-of-Thought (Sci-RCoT) to explicitly model underlying visual logic. For evaluation, we propose SciIR-Bench, which aligns with these three semiotic levels and employs an Atomic Checklist to convert the outcome-oriented scientific accuracy into process-oriented, verifiable, fine-grained questions. Our extensive experiments reveal significant deficiencies in current models' scientific reasoning capabilities. Furthermore, by fine-tuning on the SciIR-82k dataset, we developed the Qwen-Image-SciIR model, which achieves a substantial improvement on the SciIR-Bench, increasing the final score from 35\% to 43\%, laying a solid foundation for future advances in scientific image generation.

科学图像符号学图像生成推理链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。