顶尖大模型在欧洲高管决策任务上表现远低于专家水平。
EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks
- 用413个真实案例构建专家评测基准,人工评估生成答案质量。
- 最强模型仅解决56.9%任务,专家参考答案接近全对且更受青睐。
- 揭示自动评估在主观性强的真实问题上存在明显不足,适合政策与管理研究者参考。
前沿大语言模型正被用于开放式的复杂问题,这类问题与常规评估任务性质不同。我们投入超过4000小时专家人力,评估六款前沿大模型在新提出的EuroExec基准上的表现——该基准包含413个由47位经认证领域专家基于真实案例撰写的开放式长文本欧洲高管决策任务。每个回答均通过多维度评分表、项目特定检查清单及偏好排序进行人工评估,得出综合指标“求解率”。最强模型仅解决56.9%的任务,而专家参考答案在盲评中几乎达到天花板水平,且在74%的直接对比中胜过所有模型输出。结果表明,当前生成系统在专业级任务中仍显著落后于人类专家。我们发现,唯有依赖人工评估并结合严格统计分析确保一致性,才能获得可靠结论;同时自动测评方法在具有主观真实性的开放问题上同样表现不佳。
原文摘要 · Abstract (English)
Frontier LLMs are increasingly put to use on open-ended complex questions, different in nature from the ones they are typically evaluated on. We dedicate more than 4,000 human expert hours to evaluate a selection of six frontier LLMs on a member of this class of problems: EuroExec, our introduced human expert-based benchmark composed of 413 open-ended long-form European executive tasks authored by 47 vetted domain experts, each question drawn from experience in a real case. Every response is manually evaluated through a multi-attribute rubric, an item-specific checklist of requirements, and a preference rank ordering, extracting an aggregate metric "Solve Rate". The strongest model solves only 56.9% of tasks, while expert-written reference answers judged blindly are solved at near-ceiling levels and are preferred over every model response in 74% of direct rankings, placing frontier generative systems well below the professional standard of work they are already used for. We see that the best way to extract this kind of conclusion is by employing human evaluators, carefully checking their consistency through rigorous statistical analysis, and observe that automatic measurements also fall short when evaluating on this case of real-world open-ended problems with a subjective ground truth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。