用随机变量题测试大模型真推理能力,发现其泛化有限但可提升。
Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions
- 构建生成函数生成带随机变量的未见题目
- 30+模型在1000+题上测试,准确率随变量变化波动大
- 揭示模型在新题型上泛化弱,但可通过测试时扩展改善
现有数学评测存在设计简单和数据污染等问题,难以真实评估大语言模型(LLMs)的数学推理能力。为此,我们提出RV-Bench,一种基于随机变量的新型评测方法。通过构建问题生成函数,生成背景与原题相似但变量组合随机化的未见题目(RVQs),迫使模型理解内在模式而非记忆答案。我们在超过30个代表性LLMs上进行了大规模实验,覆盖1000多道RVQ。结果表明,模型在已见数据上表现良好,但在未见数据分布下出现能力不平衡;尽管在相似任务间泛化能力有限,但通过测试时缩放(test-time scaling)可有效激发其推理潜力。
原文摘要 · Abstract (English)
Recent studies have raised significant concerns regarding the reliability of current mathematics benchmarks, highlighting issues such as simplistic design and potential data contamination. Consequently, developing a reliable benchmark that effectively evaluates large language models' (LLMs) genuine capabilities in mathematical reasoning remains a critical challenge. To address these concerns, we propose RV-Bench, a novel evaluation methodology for Benchmarking LLMs with Random Variables in mathematical reasoning. Specifically, we build question-generating functions to produce random variable questions (RVQs), whose background content mirrors original benchmark problems, but with randomized variable combinations, rendering them "unseen" to LLMs. Models must completely understand the inherent question pattern to correctly answer RVQs with diverse variable combinations. Thus, an LLM's genuine reasoning capability is reflected through its accuracy and robustness on RV-Bench. We conducted extensive experiments on over 30 representative LLMs across more than 1,000 RVQs. Our findings propose that LLMs exhibit a proficiency imbalance between encountered and ``unseen'' data distributions. Furthermore, RV-Bench reveals that proficiency generalization across similar mathematical reasoning tasks is limited, but we verified it can still be effectively elicited through test-time scaling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。