arXiv:2502.07813cs.CRcs.AI2025-02被引 5

首个量化大模型组合推理能力的评估框架

CryptoX : Compositional Reasoning Evaluation of Large Language Models

  • 融合密码学与现有基准,构建可量化组合推理的新方法
  • 发现开源与闭源大模型在组合推理上存在巨大差距
  • 揭示模型分解问题、推理子问题、总结结论的内在机制

组合推理能力长期被视为大语言模型(LLMs)泛化与智能涌现的关键。然而,现有众多推理类基准极少对组合推理能力进行系统研究或量化评估。本文提出CryptoX,首个结合密码学原理与现有基准的评估框架,用于量化LLMs的组合推理能力。基于CryptoX,我们构建了CryptoBench,将该原则融入多个基准以实现系统性评估。我们在广泛使用的开源与闭源大模型上开展实验,揭示了开源与闭源模型在组合推理能力上的显著差距。进一步通过机械可解释性实验,分析了模型在子问题分解、子问题推理及结论归纳方面的内部机制。基于CryptoBench的分析强调了独立研究组合推理的价值,并呼吁增强大模型的组合推理能力。

原文摘要 · Abstract (English)

The compositional reasoning capacity has long been regarded as critical to the generalization and intelligence emergence of large language models LLMs. However, despite numerous reasoning-related benchmarks, the compositional reasoning capacity of LLMs is rarely studied or quantified in the existing benchmarks. In this paper, we introduce CryptoX, an evaluation framework that, for the first time, combines existing benchmarks and cryptographic, to quantify the compositional reasoning capacity of LLMs. Building upon CryptoX, we construct CryptoBench, which integrates these principles into several benchmarks for systematic evaluation. We conduct detailed experiments on widely used open-source and closed-source LLMs using CryptoBench, revealing a huge gap between open-source and closed-source LLMs. We further conduct thorough mechanical interpretability experiments to reveal the inner mechanism of LLMs' compositional reasoning, involving subproblem decomposition, subproblem inference, and summarizing subproblem conclusions. Through analysis based on CryptoBench, we highlight the value of independently studying compositional reasoning and emphasize the need to enhance the compositional reasoning capabilities of LLMs.

组合推理评估框架大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。