3.8B小模型通过定向推理实现与GPT-4o相当的事实准确率。
Humains-Junior: A 3.8B Language Model Achieving GPT-4o-Level Factual Accuracy by Directed Exoskeleton Reasoning
- 用最小化定向'外骨骼推理'框架结合行为微调,提升模型事实判断能力。
- 在Q1-Q500测试中,准确率达72.7%,与GPT-4o差异仅0.8个百分点(±5pp内等效)。
- 成本仅为GPT-4o的1/19,适合边缘部署和低成本应用。
我们提出Humans-Junior,一个3.8B参数模型,在相同评估者下于FACTS Grounding公开子集上达到与GPT-4o相当的性能(±5个百分点等效范围内)。在Q1–Q500测试中,GPT-4o得分为73.5%(95%置信区间69.5–77.2),Humans-Junior为72.7%(95%置信区间68.7–76.5);配对差异为0.8个百分点(置换检验p=0.72,Cohen's d=0.023)。TOST检验确认在±5个百分点内等效,但不满足±3个百分点。以管理API形式部署时,Humans-Junior基础模型(Phi-3.5-mini-instruct)在Microsoft AI Foundry定价下约为GPT-4o的1/19;自托管或边缘部署可使推理成本趋近于零。方法上,结合最小化定向“外骨骼推理”结构与行为微调,训练模型遵守认知规范(知识边界意识),而非直接学习答案。单独使用微调效果有限,二者结合后显著提升准确率(+17.7个百分点,p<0.001)并降低方差约25%。在前沿模型提示仅设置中(Q1–Q100;不可比),定向推理使GPT-4o提升+11.8个百分点至85.3%,Gemini-2.5-Pro提升+5.0个百分点至93.3%(基线88.3%,n=100)。定价来源详见附录E。探索性结果见附录F。
原文摘要 · Abstract (English)
We introduce Humans-Junior, a 3.8B model that matches GPT-4o on the FACTS Grounding public subset within a $\pm 5$ pp equivalence margin. Results. On Q1--Q500 under identical judges, GPT-4o scores 73.5% (95% CI 69.5--77.2) and Humans-Junior 72.7% (95% CI 68.7--76.5); the paired difference is 0.8 pp (bootstrap 95% CI $-3.1$ to $+4.7$; permutation $p = 0.72$; Cohen's $d = 0.023$). TOST establishes equivalence at $\pm 5$ pp (not at $\pm 3$ pp). When purchased as managed APIs, Humans-Junior's base model (Phi-3.5-mini-instruct) is $\approx 19\times$ less expensive than GPT-4o on Microsoft AI Foundry pricing; self-hosted or edge deployments can drive incremental inference cost toward zero. Measured vs estimated pricing sources are tabulated in Appendix E. Method. Our approach combines minimal directed "Exoskeleton Reasoning" scaffolds with behavioral fine-tuning that teaches protocol compliance (epistemic discipline) rather than domain answers. Fine-tuning alone adds little; combined, they synergize (+17.7 pp, $p < 0.001$) and reduce variance ($\approx 25\%$). In prompt-only settings on frontier models (Q1--Q100; non-comparable), directed reasoning improved GPT-4o by +11.8 pp to 85.3% and Gemini-2.5-Pro by +5.0 pp to 93.3% (baseline 88.3%, $n = 100$); see Section~5. TL;DR. A 3.8B model achieves GPT-4o-level FACTS accuracy (equivalent within $\pm 5$ pp on Q1--Q500). Cloud pricing shows $\approx 19\times$ lower cost versus GPT-4o, and self-hosted/edge deployments can approach zero marginal cost. Pricing sources are listed in Appendix E. Frontier prompt-only gains (Q1--Q100; non-comparable) and optimized-prompt exploratory results under earlier judges are summarized in Appendix F. Keywords: Small Language Models, Factual Grounding, Directed Reasoning, Fine-Tuning, Model Alignment, Cost-Efficient AI
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。