构建葡萄酒领域评测基准,验证大模型知识推理能力。
OenoBench: A Wine-Domain Benchmark for Knowledge-Grounded Evaluation of Large Language Models

- 用16个主流模型在3266道题上测试,基于真实来源的权威数据。
- 最高准确率83.6%(o3模型),开放域问答提升33个百分点。
- 揭示模型对自身问题偏好差异,适合评估模型可靠性与知识边界。
我们提出OenoBench,一个包含3,266道多项选择题的葡萄酒领域知识评测基准,覆盖六大维度(产区、葡萄品种、栽培、酿造、酒庄、商业)和四个难度层级。数据源自38,104条经验证的原子事实,来自政府注册机构(INAO、TTB、OIV)、同行评审期刊及维基百科/Wikidata,由35个溯源可靠的爬虫采集。方法上采用大模型驱动的流水线:语言模型负责格式化与审计,但不作为真理来源;每条主张均附带原始链接,每道题目由五种策略生成,涵盖五类生成器,最终通过九名代理组成的审计系统校准,依据人类黄金标注集计算Cohen's κ。评估16个前沿模型配置后发现:(i) 整体准确率53%-84%,o3达83.6%;(ii) DeepSeek R1在推理模式下提升6.8个百分点,而Claude Opus与Gemini Pro无提升;(iii) Anthropic模型对其自身问题偏好+9个百分点,谷歌模型则呈现-8个百分点逆向偏好;(iv) 开源前沿模型与专有推理模型共享成本-精度帕累托前沿;(v) 所有配置在闭卷可解题上平均提升33个百分点,揭示参数化召回存在上限,仅上下文信息能突破此限制。我们公开数据集、审计结果、人工审核工具及构建代码,许可协议为CC-BY-SA-4.0。
原文摘要 · Abstract (English)
We introduce OenoBench, a wine-domain knowledge benchmark of 3,266 multiple-choice questions across six pillars (regions, grape varieties, viticulture, winemaking, producers, business) and four difficulty tiers. The corpus is built from 38,104 atomic, source-anchored facts extracted by 35 provenance-verified scrapers from government registries (INAO, TTB, OIV), peer-reviewed journals, and Wikipedia/Wikidata. Our methodological contribution is an LLM-driven pipeline in which language models reformat verified facts and audit the result, but never serve as the source of truth: every claim traces to a URL, every question is generated by one of five strategies across five generator families, and every question is scored by a nine-agent audit calibrated against a human gold sheet via Cohen's $κ$. Evaluating sixteen frontier configurations, we find: (i) overall accuracy spans 53%-84%, led by o3 at 83.6%; (ii) reasoning-mode lift concentrates in DeepSeek R1 (+6.8pp) and is absent in Claude Opus and Gemini Pro; (iii) Anthropic shows +9pp self preference on its own questions while Google shows -8pp inverse preference; (iv) frontier open-weight models share the cost-vs-accuracy Pareto frontier with proprietary reasoning models; and (v) every config gains around 33pp on closed-book solvable items, revealing a parametric-recall ceiling that only the contextual slice avoids. We release corpus, audit findings, human-review app, and construction code under CC-BY-SA-4.0.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。