arXiv:2601.17027cs.CVcs.AI2026-01被引 7

用逻辑驱动框架生成科学图像,提升下游推理能力。

Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility

  • 提出ImgCoder框架,分三步理解、规划、编码生成更精准图像。
  • 发现像素生成模型普遍存在科学错误,存在表达力与精度的权衡。
  • 在严谨合成图像上微调多模态模型,推理性能显著提升。

尽管合成数据在文本领域已证明能提升科学推理能力,但多模态推理仍受限于科学图像生成的困难。现有文本到图像(T2I)模型生成的图像虽视觉可信,却常缺乏科学正确性,导致视觉与逻辑的持续偏离,限制其下游推理价值。受新一代T2I模型启发,我们系统研究了不同生成范式、评估方法与下游应用。分析了直接像素生成与程序化合成两种方式,提出ImgCoder框架,采用显式的“理解-规划-编码”流程以增强结构精确性。为严格评估科学正确性,引入SciGenBench,基于信息效用与逻辑有效性进行评测。结果揭示像素生成模型存在系统性失败模式,并凸显表达力与精度之间的根本权衡。最终表明,在经严格验证的合成科学图像上微调大型多模态模型(LMMs),可实现稳定的推理提升,且具备类似文本领域的扩展趋势,验证了高保真科学图像合成作为解锁大规模多模态推理能力的可行路径。

原文摘要 · Abstract (English)

While synthetic data has proven effective for improving scientific reasoning in the text domain, multimodal reasoning remains constrained by the difficulty of synthesizing scientifically rigorous images. Existing Text-to-Image (T2I) models often produce outputs that are visually plausible yet scientifically incorrect, resulting in a persistent visual-logic divergence that limits their value for downstream reasoning. Motivated by recent advances in next-generation T2I models, we conduct a systematic study of scientific image synthesis across generation paradigms, evaluation, and downstream use. We analyze both direct pixel-based generation and programmatic synthesis, and propose ImgCoder, a logic-driven framework that follows an explicit "understand - plan - code" workflow to improve structural precision. To rigorously assess scientific correctness, we introduce SciGenBench, which evaluates generated images based on information utility and logical validity. Our evaluation reveals systematic failure modes in pixel-based models and highlights a fundamental expressiveness-precision trade-off. Finally, we show that fine-tuning Large Multimodal Models (LMMs) on rigorously verified synthetic scientific images yields consistent reasoning gains, with potential scaling trends analogous to the text domain, validating high-fidelity scientific synthesis as a viable path to unlocking massive multimodal reasoning capabilities.

图像生成科学推理多模态评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。