让图像生成模型学会科学推理,不再拍脑门造图。
Science-T2I: Addressing Scientific Illusions in Image Synthesis
- 构建2万+对抗性图文对,测试模型对科学线索的理解能力。
- 现有模型在隐含科学提示下平均得分不足50,显式提示下提升35分。
- 提出新评分模型SciScore和两阶段对齐框架,显著提升科学准确性。
当前图像生成模型虽视觉逼真,却常产生不符合物理规律的图像,暴露出视觉保真度与物理真实性之间的根本差距。本文提出ScienceT2I,一个专家标注的数据集,包含超过20,000组对抗性图像对和9,000个跨16个科学领域的提示词,以及一个包含454个高难度提示的独立测试集。基于该基准,评估18个主流图像生成模型发现,其在隐含科学提示下的平均得分均未超过50(满分100),而显式提示下得分高出约35分,表明模型能按指令生成正确场景,但无法从科学线索推断出正确视觉结果。为此,我们开发了SciScore——一个基于CLIP-H微调的奖励模型,可捕捉精细科学现象,无需依赖语言推理,在性能上超越GPT-4o和经验人类评估者约5分。进一步提出两阶段对齐框架,结合监督微调与掩码在线微调,将科学知识注入生成模型。在FLUX.1[dev]上的实验显示,该方法使SciScore得分相对提升超50%,证明通过针对性数据与对齐策略可大幅改善图像生成中的科学推理能力。
原文摘要 · Abstract (English)
Current image generation models produce visually compelling but scientifically implausible images, exposing a fundamental gap between visual fidelity and physical realism. In this work, we introduce ScienceT2I, an expert-annotated dataset comprising a training set of over 20k adversarial image pairs and 9k prompts across 16 scientific domains and an isolated test set of 454 challenging prompts. Using this benchmark, we evaluate 18 recent image generation models and find that none scores above 50 out of 100 under implicit scientific prompts, while explicit prompts that directly describe the intended outcome yield scores roughly 35 points higher, confirming that current models can render correct scenes when told what to depict but cannot reason from scientific cues to the correct visual outcome. To address this, we develop SciScore, a reward model fine-tuned from CLIP-H that captures fine-grained scientific phenomena without relying on language-guided inference, surpassing GPT-4o and experienced human evaluators by roughly 5 points. We further propose a two-stage alignment framework combining supervised fine-tuning with masked online fine-tuning to inject scientific knowledge into generative models. Applying this framework to FLUX.1[dev] yields a relative improvement exceeding 50% on SciScore, demonstrating that scientific reasoning in image generation can be substantially improved through targeted data and alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。