动态生成分布外数据,更真实评估大模型推理鲁棒性
ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning
- 动态构建分布外数据集,避免答案泄露问题
- 16个大模型测试显示多数表现不稳健,存在数据泄露
- 适合评估模型真实推理能力,尤其关注泛化性能的研究者
评估大语言模型(LLMs)面临数据污染和正确答案泄露等挑战。为此,我们提出ThinkBench,一种新型评估框架,用于可靠评估LLMs的推理能力。ThinkBench采用动态数据生成方法构建分布外(OOD)数据集,并提供包含2,912个样本的OOD数据集,涵盖多种推理任务。该框架统一评估推理模型与非推理模型。我们在相同实验条件下评估了16个LLMs和4个PRMs,发现大多数LLMs的表现远未达到稳健,且存在一定程度的数据泄露。通过动态生成OOD数据集,ThinkBench有效实现可靠评估,显著降低数据污染影响。
原文摘要 · Abstract (English)
Evaluating large language models (LLMs) poses significant challenges, particularly due to issues of data contamination and the leakage of correct answers. To address these challenges, we introduce ThinkBench, a novel evaluation framework designed to evaluate LLMs' reasoning capability robustly. ThinkBench proposes a dynamic data generation method for constructing out-of-distribution (OOD) datasets and offers an OOD dataset that contains 2,912 samples drawn from reasoning tasks. ThinkBench unifies the evaluation of reasoning models and non-reasoning models. We evaluate 16 LLMs and 4 PRMs under identical experimental conditions and show that most of the LLMs' performance are far from robust and they face a certain level of data leakage. By dynamically generating OOD datasets, ThinkBench effectively provides a reliable evaluation of LLMs and reduces the impact of data contamination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。