测试大模型在复杂推理中的泛化与鲁棒性,发现推理模型表现更好但难应对不确定性。
I-RAVEN-X: Benchmarking Generalization and Robustness of Analogical and Mathematical Reasoning in Large Language and Reasoning Models
- 构建更复杂的符号推理基准,提升运算复杂度和感知不确定性
- 推理模型在长链推理和宽范围属性上表现优于语言模型
- 模型仍难以处理不确定情境下的多概率结果探索
我们提出 I-RAVEN-X,一个用于评估大语言模型(LLMs)和大推理模型(LRMs)在类比与数学推理中泛化能力与鲁棒性的符号基准。该基准在 I-RAVEN 基础上提升了操作数复杂度、属性取值范围,并引入感知不确定性。实验表明,相较于 LLMs,LRMs 在更长的推理链上表现出更高的效率,在更广的属性范围内展现出更强的系统性。然而,LRMs 在不确定性推理下仍面临显著挑战,无法有效探索多种可能的结果分布。
原文摘要 · Abstract (English)
We introduce I-RAVEN-X, a symbolic benchmark designed to evaluate generalization and robustness in analogical and mathematical reasoning for Large Language Models (LLMs) and Large Reasoning Models (LRMs). I-RAVEN-X extends I-RAVEN by increasing operand complexity, attribute range, and introducing perceptual uncertainty. Compared to LLMs, empirical results show that LRMs achieve improved productivity and systematicity on longer reasoning relations and wider attribute ranges, respectively. However, LRMs are still significantly challenged by reasoning under uncertainty and cannot effectively explore multiple probabilistic outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。