用统计方法让大模型自修正科学推理,大幅减少错误。
Though Language Models Err While They Strive: Conformal Prediction for Self-Correcting Scientific Generation

- 构建图结构校验框架,逐步验证科学事实与逻辑连贯性。
- 在物理推理任务上达50.1%准确率,错误率降低73%。
- 适合需要高可信科学生成的科研与教育场景。
大语言模型在生成技术内容时常违反基本科学原理,影响其在科研中的可靠性。本文提出科学可行性控制(SFC),一种基于图结构的置信区间预测框架,通过逐步进行绝对-连贯-真实性验证,为科学推理提供统计保障。该方法将科学推理分解为原子级的绝对-连贯-真实性单元,要求每个单元既符合物理定律,又能在前序上下文中得到逻辑支持,从而缓解早期错误导致的连锁错误问题。不同于独立假设的方法,SFC将逻辑依赖建模为近似演绎图,在检测到科学违规时动态分支,利用已验证上下文引导替代生成路径。我们在PhyX多模态物理、MATH、ScienceQA和ARC Challenge等基准上验证SFC,取得50.1%的物理推理准确率,显著优于DeepSeek-R1(49.8%)和GPT-4(45.8%),同时在α=0.10置信水平下实现91.7%的科学有效性覆盖率,跨多种模型架构减少73%的科学法则违反。
原文摘要 · Abstract (English)
Large language models frequently violate fundamental scientific principles when generating technical content, undermining their reliability in scientific applications. We introduce Scientific Feasibility Control SFC, a graph-structured conformal prediction framework that provides statistical guarantees for scientific reasoning validity through progressive absolute-coherent-factuality validation. Our approach decomposes scientific reasoning into atomic absolute-coherent-factuality units requiring both individual correctness against physical laws and logical substantiation from preceding context, addressing the cascade effect where early scientific errors contaminate subsequent reasoning steps. Unlike independence-based methods that treat claims in isolation, SFC models logical dependencies as approximate deducibility graphs and operates through real-time validation with dynamic branching when scientific violations are detected, the system branches to alternative generation paths using verified context as foundation. We demonstrate SFC across established scientific reasoning benchmarks including PhyX multimodal physics, MATH, ScienceQA, and ARC Challenge, achieving 50.1 percent accuracy on PhyX physics reasoning, substantially outperforming recent reasoning models including DeepSeek-R1 49.8 percent and GPT-4 45.8 percent while providing 91.7 percent scientific validity with formal conformal coverage guarantees at alpha equals 0.10 confidence level and reducing scientific law violations by 73 percent across multiple model architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。