arXiv:2608.13112cs.CV2026-08

让生成的物理图示真正符合科学原理,而非仅看起来像。

Towards Physics-Faithful Generation of Scientific Diagrams

论文配图:Towards Physics-Faithful Generation of Scientific Diagrams
图 1 · 摘自论文原文
  • 用结构化物理思维链分解图像生成过程,确保每一步推理有据可循。
  • 构建430万张带结构标注的物理图示数据集,11.5万张由专家标注。
  • 提出专用评测基准,可逐项检验生成图的物理正确性,适合教育与科研场景。

文本生成图像已达到逼真水平,但当前主流系统在生成科学图表时仍不可靠,其价值不在于外观而在于物理正确性:力的方向、坐标系的有效性、热力学状态的一致性以及方程与场景的匹配。现有模型基于网络图像和浅层描述训练,生成的图表看似合理实则物理错误,影响教学与科学传播。本文提出普林西格拉姆(Princigram)及其数据管道,核心创新为结构化物理思维链(SP-CoT):按学科细分,将物理图示生成分解为从场景识别到受力分析、规律应用及整合的六步结构化推理链。该链路遵循固定模式,严格区分视觉事实与物理推断,并对数学符号进行类型标注,既作为密集监督信号,又在推理时充当结构化“思考”提示。我们据此构建并结构化标注了430万张物理图像,其中115,037张为专家级标注,并适配统一多模态主干模型。进一步提出VeriphyT2IBench评测基准,其问题源自每个保留图示的结构化标注:每张图对应一组关于物体、力与状态的二元问题,使评估结果可分解为具体物理事实,而非单一综合评分。在GenExam物理子集及VeriphyT2IBench上,Princigram验证了显式物理结构化监督显著提升生成图表的物理忠实度。

原文摘要 · Abstract (English)

Text-to-image generation has reached photorealistic quality, yet state-of-the-art systems remain unreliable at producing scientific diagrams, whose value depends not on appearance but on physical faithfulness: correct force directions, valid coordinate systems, consistent thermodynamic states, and equations matching the depicted scenario. Trained on web imagery with physically shallow captions, generic models produce diagrams that look plausible but are physically wrong, harmful in education and scientific communication. We present Princigram, a physics-faithful scientific-diagram generator, and its data pipeline. Our central advance is Structured Physical Chain-of-Thought (SP-CoT): a per-subdiscipline schema that decomposes a physics diagram into an explicit multi-step reasoning chain across six subdisciplines, from scene identification through force or process analysis to governing laws and synthesis. Unlike free-form chain-of-thought, SP-CoT follows a fixed schema with strict fidelity rules that separate visually grounded facts from physically inferred reasoning and type all mathematics symbolically; it serves both as dense training supervision and, at inference, as a structured "thinking" prompt. With it we curate and structurally annotate 4.3 million physics images, of which 115,037 carry expert-level annotation, and adapt a unified multimodal backbone. We further introduce VeriphyT2IBench, whose questions are derived from each held-out diagram's own structured annotation: each diagram becomes an item-specific bank of binary questions about its objects, forces, and states, so a judge model's score decomposes into named physical facts rather than one holistic number. On the physics subset of GenExam and on VeriphyT2IBench, Princigram shows that explicit physics-structured supervision improves the physical faithfulness of generated scientific diagrams.

科学绘图物理生成结构化推理图文生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。