用商学院案例测试AI的分析决策能力,发现顶尖模型已表现优异且持续进步。
Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

- 基于商学院真实案例构建评测基准,覆盖十八个学科领域。
- 前沿AI模型在该基准上得分高,同一模型家族两年间性能显著提升。
- 适合关注AI在商业分析、战略推理中应用的研究者与教育从业者。
大型语言模型(LLMs)在事实记忆、数学求解、编程等任务上的表现迅速提升,但对白领专业人士日常从事的分析性知识工作——如整合复杂信息、在不确定和不完整数据下做出判断、多利益相关方情境下的策略与对抗性思维、权衡取舍并生成可辩护的结构化分析——的评估仍不足,尤其主观性较强的工作更难量化。本文借鉴顶尖商学院的案例教学法,构建了涵盖十八个学科领域、数百道题目的BusinessCaseBench基准,每道题均配有依据专家撰写教案制定的评分标准。实验显示,当前前沿AI模型在该基准上表现良好,且同一模型家族在两年内能力大幅提升。结果表明,AI在这一类分析性工作中已具备较高水平并快速演进,对商学院教学及初级职业岗位的技能要求产生深远影响。
原文摘要 · Abstract (English)
Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use. What remains poorly measured is AI progress on the analytical knowledge work white-collar professionals perform daily, including synthesizing complex information, exercising judgment under uncertainty and incomplete information, applying strategic and adversarial thinking in multi-stakeholder settings, weighing trade-offs, and producing defensible, structured analyses. This gap is even more pronounced for subjective components of such work, where success can be challenging to define. The "case method" form of education practiced by top business schools provides a natural foundation for addressing this measurement gap, and we construct BusinessCaseBench, a benchmark spanning hundreds of questions drawn from business cases across eighteen disciplines, each paired with a grading rubric derived from the expert-written instructor case solution. On BusinessCaseBench, frontier AI models already score highly against instructor rubrics, and capability within one model family improves substantially over two years. These results provide strong evidence that AI performance on this class of work is already high and rapidly improving, with implications for business schools, where case pedagogy trains undergraduates and MBAs in this kind of analytical reasoning, and for entry-level professional roles, where such skills have historically anchored early-career work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。