评测大模型在20+理工科领域的解题能力,发现其推理链稳定性不足。
Classroom Final Exam: An Instructor-Tested Reasoning Benchmark
- 基于真实大学考试题构建多模态推理基准
- 顶尖模型平均准确率仅59.7%,仍有巨大提升空间
- 模型推理步数多但易出错,不如人类教师解题高效
我们提出CFE-Bench(课堂期末考),一个涵盖20多个理工科领域的多模态推理评估基准。该数据集源自多次使用的大学作业与考试真题,并配有授课教师提供的参考解答。当前前沿模型表现仍不理想:新发布的Gemini-3.1-pro-preview整体准确率为59.69%,第二好的Gemini-3-flash-preview为55.46%,改进空间显著。通过将教师参考解答分解为结构化推理流程,我们发现尽管模型常能正确回答中间子问题,但在多步推理中难以稳定维持正确的中间状态。此外,模型生成的解题步骤普遍多于教师解答,表明其推理效率低,错误累积风险高。数据与代码已开源:https://github.com/Analogy-AI/CFE_Bench。
原文摘要 · Abstract (English)
We introduce CFE-Bench (Classroom Final Exam), a multimodal benchmark for evaluating the reasoning capabilities of large language models across more than 20 STEM domains. CFE-Bench is curated from repeatedly used, authentic university homework and exam problems, paired with reference solutions provided by course instructors. CFE-Bench remains challenging for frontier models: the newly released Gemini-3.1-pro-preview achieves 59.69% overall accuracy, while the second-best model, Gemini-3-flash-preview, reaches 55.46%, leaving substantial room for improvement. Beyond aggregate scores, we conduct a diagnostic analysis by decomposing instructor reference solutions into structured reasoning flows. We find that while frontier models often answer intermediate sub-questions correctly, they struggle to reliably derive and maintain correct intermediate states throughout multi-step solutions. We further observe that model-generated solutions typically contain more reasoning steps than instructor solutions, indicating lower step efficiency and a higher risk of error accumulation. Data and code are available at https://github.com/Analogy-AI/CFE_Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。