arXiv:2604.19758cs.AIcs.CL2026-04

评测大模型在热力学推理上的能力,分三层难题挑战真实工程分析能力。

ThermoQA: A Three-Tier Benchmark for Evaluating Thermodynamic Reasoning in Large Language Models

论文配图:ThermoQA: A Three-Tier Benchmark for Evaluating Thermodynamic Reasoning in Large Language Models
图 1 · 摘自论文原文
  • 构建三层次热力学问题集,涵盖查表、部件分析到完整循环计算。
  • 顶尖模型正确率超92%,但跨层级性能下降达32.5个百分点,暴露推理短板。
  • 聚焦超临界水等复杂场景,揭示模型真实热力理解差异,适合评估工程级AI。

我们提出ThermoQA,一个包含293道开放性工程热力学问题的基准测试,分为三层:属性查询(110题)、组件分析(101题)和完整循环分析(82题)。真实答案通过CoolProp 7.2.0程序化计算得出,覆盖水、R-134a及变比热空气。六种前沿大模型在三个独立运行中被评估。综合排行榜由Claude Opus 4.6(94.1%)、GPT-5.4(93.1%)和Gemini 3.1 Pro(92.5%)领先。跨层级性能下降从2.8个百分点(Opus)到32.5个百分点(MiniMax),证明属性记忆不等于热力学推理能力。超临界水、R-134a制冷剂与联合循环燃气轮机分析作为自然区分器,性能差距达40-60个百分点。多轮次标准差范围为±0.1%至±2.5%,量化了推理一致性这一新评估维度。数据集与代码开源于https://huggingface.co/datasets/olivenet/thermoqa。

原文摘要 · Abstract (English)

We present ThermoQA, a benchmark of 293 open-ended engineering thermodynamics problems in three tiers: property lookups (110 Q), component analysis (101 Q), and full cycle analysis (82 Q). Ground truth is computed programmatically from CoolProp 7.2.0, covering water, R-134a, and variable-cp air. Six frontier LLMs are evaluated across three independent runs each. The composite leaderboard is led by Claude Opus 4.6 (94.1%), GPT-5.4 (93.1%), and Gemini 3.1 Pro (92.5%). Cross-tier degradation ranges from 2.8 pp (Opus) to 32.5 pp (MiniMax), confirming that property memorization does not imply thermodynamic reasoning. Supercritical water, R-134a refrigerant, and combined-cycle gas turbine analysis serve as natural discriminators with 40-60 pp performance spreads. Multi-run sigma ranges from +/-0.1% to +/-2.5%, quantifying reasoning consistency as a distinct evaluation axis. Dataset and code are open-source at https://huggingface.co/datasets/olivenet/thermoqa

热力学大模型评测工程推理基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。