打造一个挑战性极强的机器学习知识评测基准,专测AI对数据科学的理解与推理能力。
HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI
- 构建100道原创多选题,覆盖现代机器学习核心内容
- 当前顶尖AI模型在该基准上错误率高达30%,远超主流基准
- 适合评估高阶AI在数据科学领域的认知水平,填补领域空白
我们提出HardML,一个用于评估人工智能在数据科学与机器学习领域知识和推理能力的基准。HardML包含100道精心设计的难题,历时6个月手工制作,涵盖数据科学与机器学习中最新且最热门的分支。这些题目对普通高级机器学习工程师也极具挑战性。为降低数据污染风险,大多数内容为作者原创。当前顶级AI模型在此基准上的错误率达30%,约为知名基准MMLU ML的3倍。尽管受限于多选题形式,无法推动前沿发展,但HardML仍可作为严谨、现代化的测试平台,用于量化和追踪顶尖AI的进步。目前,数学、物理、化学等领域已有大量LLM评测基准,而数据科学与机器学习子领域仍严重缺乏此类评估体系。
原文摘要 · Abstract (English)
We present HardML, a benchmark designed to evaluate the knowledge and reasoning abilities in the fields of data science and machine learning. HardML comprises a diverse set of 100 challenging multiple-choice questions, handcrafted over a period of 6 months, covering the most popular and modern branches of data science and machine learning. These questions are challenging even for a typical Senior Machine Learning Engineer to answer correctly. To minimize the risk of data contamination, HardML uses mostly original content devised by the author. Current state of the art AI models achieve a 30% error rate on this benchmark, which is about 3 times larger than the one achieved on the equivalent, well known MMLU ML. While HardML is limited in scope and not aiming to push the frontier, primarily due to its multiple choice nature, it serves as a rigorous and modern testbed to quantify and track the progress of top AI. While plenty benchmarks and experimentation in LLM evaluation exist in other STEM fields like mathematics, physics and chemistry, the subfields of data science and machine learning remain fairly underexplored.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。