arXiv:2409.18486cs.CL2024-09被引 172

o1-preview在多领域复杂推理任务中表现接近或超越人类,展现强通用智能潜力。

Evaluation of OpenAI o1: Opportunities and Challenges of AGI

  • 跨领域测试验证其在编程、医学、数学等任务中的综合推理能力。
  • 数学题全对、编程正确率达83.3%,医学报告生成优于其他模型。
  • 适合关注AGI进展的研究者与开发者,尤其关注通用推理应用。

本研究全面评估了OpenAI o1-preview大语言模型在计算机科学、数学、自然科学、医学、语言学及社会科学等多个领域的复杂推理任务表现。通过严格测试,o1-preview在诸多任务中展现出接近或超越人类的水平:在复杂编程竞赛中成功率达83.3%,显著高于多数人类专家;生成放射科报告的能力优于其他对比模型;高中数学推理任务准确率100%,并提供详细解题步骤;在通用与医学等专业领域具备先进的自然语言推理能力;芯片设计任务中,于EDA脚本生成和缺陷分析方面超越专用模型;在人类学与地质学领域表现出深度理解与推理能力;在量化投资方面具备全面金融知识与统计建模能力;社会媒体分析中,情感识别与情绪判断表现优异。尽管在部分简单问题上偶有失误,且对极专业概念存在挑战,整体结果表明其已显著逼近人工通用智能(AGI)。

原文摘要 · Abstract (English)

This comprehensive study evaluates the performance of OpenAI's o1-preview large language model across a diverse array of complex reasoning tasks, spanning multiple domains, including computer science, mathematics, natural sciences, medicine, linguistics, and social sciences. Through rigorous testing, o1-preview demonstrated remarkable capabilities, often achieving human-level or superior performance in areas ranging from coding challenges to scientific reasoning and from language processing to creative problem-solving. Key findings include: -83.3% success rate in solving complex competitive programming problems, surpassing many human experts. -Superior ability in generating coherent and accurate radiology reports, outperforming other evaluated models. -100% accuracy in high school-level mathematical reasoning tasks, providing detailed step-by-step solutions. -Advanced natural language inference capabilities across general and specialized domains like medicine. -Impressive performance in chip design tasks, outperforming specialized models in areas such as EDA script generation and bug analysis. -Remarkable proficiency in anthropology and geology, demonstrating deep understanding and reasoning in these specialized fields. -Strong capabilities in quantitative investing. O1 has comprehensive financial knowledge and statistical modeling skills. -Effective performance in social media analysis, including sentiment analysis and emotion recognition. The model excelled particularly in tasks requiring intricate reasoning and knowledge integration across various fields. While some limitations were observed, including occasional errors on simpler problems and challenges with certain highly specialized concepts, the overall results indicate significant progress towards artificial general intelligence.

通用智能推理能力多领域应用大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。