揭示材料科学假说生成中机制恢复的视觉追踪方法
Visualizing Graph-to-Answer Mechanism Recovery in Materials-Science Hypothesis Generation

- 构建图到答案的机制追踪流程,可视化模型推理路径
- 早期层(7-10)几乎无机制恢复,晚期合成层(30,36)恢复显著
- 适合科学家与开发者诊断假说生成中的逻辑断裂点
AI协作者能生成流畅的材料科学假说,但流畅性不等于科学机制的保留。本文针对经微调的Graph-PRefLexOR-8B(基于Qwen3-8B)模型,开展图到答案的机制追踪案例研究,涵盖头脑风暴、图构建、模式提取与合成四个阶段。通过语义回溯、图破坏实验、激活恢复测量及逐层逐标记网格分析,构建可视化诊断流程。在100个开放问题上,最终答案仍最贴近模型自身的结构化阶段,尤其是合成阶段。在对37个残差流检查点、嵌入输出及36个Transformer块的全面测试中发现,第7至10层的早期过渡区几乎无机制恢复,而恢复主要集中在后期合成与答案起始区域(约第30和36层)。该流程旨在帮助科学家与模型开发者识别假说生成过程中机制支持的丢失或重建节点,为后续实验规划提供可靠依据。
原文摘要 · Abstract (English)
AI co-scientists can generate fluent materials-science hypotheses, but fluency does not show that an answer preserves a scientifically meaningful mechanism. We present a graph-to-answer mechanism-tracing case study for Graph-PRefLexOR-8B, a Qwen3-8B model adapted to expose distinct stages for brainstorming, graph construction, pattern extraction, and synthesis. We organize semantic backtracking, graph corruption, activation-based recovery measurements, and layer-by-token-region grids into a visual diagnostic workflow for inspecting this pathway. Across 100 open-ended materials-science questions, final answers remain closest to the model's own structured stages, especially synthesis. Under graph corruption, a full sweep over 37 residual-stream checkpoints, the embedding output and 36 transformer blocks, shows little mechanism recovery in the earlier transition region at layers 7--10, recovery instead concentrates in late synthesis and answer-start regions around layers 30 and 36. The workflow is intended to help scientists and model developers identify where a generated hypothesis loses or regains mechanism support before it is passed to downstream experimental planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。