arXiv:2608.20574cs.AIcs.CY2026-08

用可执行的菜谱评分系统评估大模型,发现微调后性能显著提升。

FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training

论文配图:FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training
图 1 · 摘自论文原文
  • 构建可计算的菜谱奖励地图,用真实评分替代人工标注。
  • 微调后模型在534个任务上平均提分13.30,效果稳定可复现。
  • 适合关注模型评测与后训练的科研人员和工程师。

开放式语言模型评估常依赖其他模型或小规模偏好小组作为答案基准。本文提出FlavourBench,通过版本化的烹饪环境生成密集的答案评分图。每项任务要求从8个候选食材中选出3个组合,推理前由Epicure对全部56种组合进行评分。我们在534个替换、配对与约束任务上评估27个前沿模型,共获得14,418次完整模型-任务观测。通过锚点聚类自举与多重性控制的配对检验,解决了351组模型对比中的101组。Grok 4.6得分最高(65.1),但校正后未识别出唯一最优模型。排名在独立评分面板下保持一致,并在不同指标、任务筛选、家族权重及三个公开Epicure检查点下稳定。我们还开展预注册的三种子微调实验:将固定版Qwen3-0.6B在270个最优菜谱上进行LoRA SFT,相比格式与标签匹配的对照组,其在84个不重叠锚点地图上提升13.30分(95%置信区间6.52至20.29,p=0.000170),且在全部534个公开地图上复现。复制增益为11.73分(95%置信区间8.98至14.54)。主划分中,两个训练组均能解析所有响应,而格式对照组未优于基线。发布内容包括提示、完整奖励图、原始响应、训练与评估清单、统计计划、代码及离线验证器。

原文摘要 · Abstract (English)

Open-ended language-model evaluation often substitutes another model or a small preference panel for a missing answer key. We introduce FlavourBench, which instead compiles dense answer maps from a versioned culinary environment. Each task asks for a three-ingredient portfolio from eight candidates; before inference, Epicure scores all 56 portfolios. We evaluate 27 frontier endpoints on the same 534 substitution, pairing, and constraint tasks, yielding 14,418 complete model-task observations. Anchor-cluster bootstraps and multiplicity-controlled paired tests resolve 101 of 351 model contrasts. Grok 4.6 has the largest point estimate at 65.1, but the corrected evidence does not identify a unique best endpoint. The ranking replicates across independently compiled panels and remains similar under alternative metrics, task filters, family weights, and three public Epicure checkpoints. We then run a preregistered, three-seed post-training study. LoRA SFT of a pinned Qwen3-0.6B checkpoint on 270 Epicure-optimal answers improves its score on 84 anchor-disjoint maps by 13.30 points over a format- and label-matched control (95% CI 6.52 to 20.29, p = 0.000170), and the effect replicates on all 534 public maps. The replication gain is 11.73 points (95% CI 8.98 to 14.54). On the primary split, both trained arms parse every response while the format control does not improve on the base model. The release contains prompts, exhaustive reward maps, raw responses, training and evaluation manifests, statistical plans, code, and an offline verifier.

模型评测微调奖励建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。