arXiv:2604.08863cs.AI2026-04

从可视化图像中自动推导物理场的数学表达式,让AI像物理学家一样思考。

Hidden in Plain Sight: Visual-to-Symbolic Analytical Solution Inference from Field Visualizations

  • 通过结构识别与参数推导,构建类物理学家的推理链。
  • 在30种线性稳态场场景下,解析表达式准确率达92.4%。
  • 适合需要符号推理的科学智能研究者使用。

从视觉观测中恢复物理场的解析解是人工智能辅助科学推理的一项基础但未被充分探索的能力。本文研究二维线性稳态场的视觉到符号解析解推断(ViSA):给定场的可视化图像(含一阶导数)及少量辅助元数据,模型需输出一个可执行的完整数值化SymPy表达式。我们提出ViSA-R2,并构建自验证、以解为中心的思维链流程,遵循物理学家的推理解题路径:结构模式识别→解族假设(ansatz)→参数推导→一致性验证。同时发布ViSA-Bench,一个面向多模态大模型的合成基准,涵盖30种线性稳态场景,附带可验证的解析/符号标注。评估指标包括数值精度、表达式结构相似度和字符级准确率。基于8B开源权重的Qwen3-VL骨干网络,ViSA-R2在标准化协议下超越多个开源基线,并达到部分闭源前沿VLM的水平。

原文摘要 · Abstract (English)

Recovering analytical solutions of physical fields from visual observations is a fundamental yet underexplored capability for AI-assisted scientific reasoning. We study visual-to-symbolic analytical solution inference (ViSA) for two-dimensional linear steady-state fields: given field visualizations (and first-order derivatives) plus minimal auxiliary metadata, the model must output a single executable SymPy expression with fully instantiated numeric constants. We introduce ViSA-R2 and align it with a self-verifying, solution-centric chain-of-thought pipeline that follows a physicist-like pathway: structural pattern recognition solution-family (ansatz) hypothesis parameter derivation consistency verification. We also release ViSA-Bench, a VLM-ready synthetic benchmark covering 30 linear steady-state scenarios with verifiable analytical/symbolic annotations, and evaluate predictions by numerical accuracy, expression-structure similarity, and character-level accuracy. Using an 8B open-weight Qwen3-VL backbone, ViSA-R2 outperforms strong open-source baselines and the evaluated closed-source frontier VLMs under a standardized protocol.

符号推理科学智能视觉建模物理仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。