首个针对高超音速热防护系统计算的诊断评估基准,专为工程安全设计。
TPS-CalcBench: A Benchmark and Diagnostic Evaluation Framework for LLM Analytical Calculation Competence in Hypersonic Thermal Protection System Engineering

- 构建4级8类任务体系,覆盖高超音速气动与高温气体动力学解析计算。
- 通过双轨评估发现13个模型准确率仅12.6%至87.9%,揭示公式误选等隐性缺陷。
- 提供微调、检索增强、过程提示三类干预方法,适合航空航天安全领域研究者。
将大语言模型应用于高安全性航空航天工程需超越通用科学评测标准。在高超音速热防护系统(TPS)设计中,驻点热流或边界层计算误差可能导致灾难性设计裕度失效。数值合理但物理无效的答案比拒答更危险。现有科学基准仅测试抽象数学与基础物理,仅评估最终答案,忽略工程推理过程,无法发现此类关键失败。本文提出TPS-CalcBench,首个面向高超音速气动与高温气体动力学闭式解析计算的诊断性基准,涵盖安德森教材中的4级8类任务;采用双轨评估机制,通过8维度评分与人工审计检测‘答案正确但推理错误’问题;构建人机协同数据流程,生成420个高置信核心题项与810个噪声控制预筛题项(来自4560条原始数据);开展噪声敏感性分析以评估数据质量对模型排名的影响;并提出三种诊断干预方法:DFA-TPS微调、RAG-EQ检索增强与PA-CoT过程感知提示。在13个来自7个团队的模型上测试显示性能差异显著(KPI 12.6–87.9),存在隐蔽公式选择缺陷,数据驱动排名变化,并验证干预有效性,建立完整的‘诊断-评估-干预’框架,用于安全关键工程中大模型部署的评估。
原文摘要 · Abstract (English)
Deploying LLMs as reasoning assistants in safety-critical aerospace engineering requires stricter evaluation criteria than general scientific benchmarks. In hypersonic thermal protection system (TPS) design, inaccurate stagnation-point heat flux or boundary-layer calculations may cause catastrophic design margin violations. Models with numerically reasonable but physically invalid answers are more dangerous than those declining to respond. Current scientific benchmarks only test abstract math and basic physics, evaluate final answers solely, ignore engineering reasoning processes, and cannot detect such critical failures. We propose TPS-CalcBench, the first diagnostic benchmark for closed-form analytical calculations in hypersonic aerodynamics and high-temperature gas dynamics that experienced TPS engineers conduct without simulations. Our contributions include domain-oriented task taxonomy with 4 difficulty levels and 8 categories from Anderson's textbook, dual-track evaluation measuring result accuracy and reasoning quality via an 8-dimension rubric and calibrated judge with human audit to identify right answer wrong reasoning issues, human-AI data pipeline producing 420 high-confidence core items and 810 noise-controlled pre-gating items from 4560 raw data, noise-sensitivity analysis measuring data quality impacts on model ranking, and three diagnostic intervention methods: DFA-TPS fine-tuning, RAG-EQ retrieval grounding and PA-CoT process-aware prompting. Tests on 13 models from 7 groups show wide performance differences (KPI 12.6-87.9), hidden formula selection defects, data-driven rank changes and effective intervention improvements, establishing a complete diagnose-evaluate-intervene framework for safety-critical engineering LLM deployment assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。