通过精确的结构力学基准,揭示了微调中标签形式如何决定模型学习内容。
OraclePhys: A Systematic Framework for LLM Fine-Tuning on Structural Mechanics
- 构建无人工标注的结构力学评测体系,用有限元模拟生成准确答案
- 发现标签格式比数据量更关键,向量标签能有效教会模型物理推理能力
- 适用于需要精准物理建模的工程场景,如建筑结构分析
现有大语言模型微调效果通常在事后评估。本文提出OraclePhys系统框架,包含三个部分:可精确定量评分的结构力学基准OraclePhys-Bench,基于字节级一致描述的七种答案形式监督数据集OraclePhys-30K,以及对七种答案形式与三种验证角色的受控训练研究。研究发现:第一,标签的输出形式(非比特数)直接影响模型所学内容——排名目标使模型具备分布外预测能力,标量目标仅获得部分能力,布尔型标签则无法被识别;该现象在第二个物理领域、第二类模型和改写评估表面均成立。第二,文本或分数筛选的答案可有效传递能力,而优势加权分数(GRPO)虽提升奖励但未显著改变模型在保留物理任务上的表现——仅够用于路由。经训练的80亿参数模型成为首个实现空间结构响应建模的通用语言模型,在零样本与32样本条件下超越基准模型,达到专家水平。模型所学内容取决于标签对目标计算的表达方式,训练内容即为模型路由依据。
原文摘要 · Abstract (English)
What a language model internalizes from fine-tuning is usually diagnosed after the fact. We make it an experimental variable. OraclePhys is a systematic fine-tuning framework with three components: OraclePhys-Bench, an exactly-graded structural-mechanics benchmark whose finite-element oracle scores every answer and counterfactual edit -- no human labels, no LLM judging; OraclePhys-30K, a supervision dataset of seven answer forms over byte-identical structure descriptions; and a controlled training study across the seven forms and three verifier roles. The study yields two findings. First, the label's answer form -- not its bit count -- causally determines what fine-tuning teaches: a ranking objective installs an out-of-distribution forward model where the untrained base sits at the guessing prior, a scalar objective at best a partial one, a boolean nothing detectable; the vector-scalar gulf survives a second physics domain, a second model family, and a paraphrased evaluation surface. Second, written or score-filtered answers install this capability, while advantage-weighted scores (GRPO) raise reward yet leave the model statistically equivalent to its start on held-out physics -- within the recipes and budgets tested -- sufficing only for routing. The trained 8B -- the first LLM on spatial structural response -- reaches the task's data-precision frontier: above a frontier LLM at zero- and 32-shot, at a specialist's level. What the label spells out about the target computation is what fine-tuning teaches; what you train on is what you route.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。