arXiv:2502.08859cs.AIcs.CL2025-02被引 21

用谜题竞赛数据集评估模型的隐性知识融合与多步推理能力

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges

  • 基于真实谜题赛事构建多模态推理挑战集
  • 顶尖模型在该基准上准确率远低于人类水平
  • 适合评估模型在非结构化场景下的创造性推理能力

随着语言模型在现有推理基准上表现趋近极限,亟需新挑战来检验其认知边界。解谜赛事蕴含大量复杂的多模态问题,能全面测试高级推理与知识整合能力,是评估前沿语言模型的理想测试场。我们提出EnigmaEval,一个源自解谜竞赛的问题与解答数据集,用于考察模型在隐性知识融合和多步演绎推理方面的能力。不同于传统基准,解谜任务要求模型发现看似无关信息间的隐藏关联,以揭示解题路径。该基准包含1184个不同复杂度的谜题,每个通常需经验丰富的团队花费数小时至数天完成,且答案明确可验证,便于高效评估。当前最先进语言模型在此基准上的准确率极低,甚至低于Humanity's Last Exam等其他困难基准,暴露出模型在面对非结构化、发散性推理任务时的显著短板。

原文摘要 · Abstract (English)

As language models master existing reasoning benchmarks, we need new challenges to evaluate their cognitive frontiers. Puzzle-solving events are rich repositories of challenging multimodal problems that test a wide range of advanced reasoning and knowledge capabilities, making them a unique testbed for evaluating frontier language models. We introduce EnigmaEval, a dataset of problems and solutions derived from puzzle competitions and events that probes models' ability to perform implicit knowledge synthesis and multi-step deductive reasoning. Unlike existing reasoning and knowledge benchmarks, puzzle solving challenges models to discover hidden connections between seemingly unrelated pieces of information to uncover solution paths. The benchmark comprises 1184 puzzles of varying complexity -- each typically requiring teams of skilled solvers hours to days to complete -- with unambiguous, verifiable solutions that enable efficient evaluation. State-of-the-art language models achieve extremely low accuracy on these puzzles, even lower than other difficult benchmarks such as Humanity's Last Exam, unveiling models' shortcomings when challenged with problems requiring unstructured and lateral reasoning.

多模态推理解谜评测认知评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。