用真实用户解谜数据评估大模型逻辑推理能力。
TurtleBench: Evaluating Top Language Models via Real-World Yes/No Puzzles
- 基于真实用户解谜行为构建动态评测集
- 9个顶级模型中o1系列未达最优表现
- 适合研究模型真实推理能力与提示工程
随着大语言模型应用的扩展,可靠评估需求日益增长。现有评估基准多依赖静态数据集,难以衡量模型在与用户动态交互中的表现;且常依赖特定背景知识,影响对逻辑推理能力的准确测量。基于强模型或人工的动态评估方法易引入偏差,成本高、耗时长,难以大规模应用。为此,我们提出TurtleBench,通过自建的Turtle Soup Puzzle平台收集真实用户猜测数据,实现相对动态的评估数据生成。该方法降低模型作弊风险,更贴近用户实际推理需求,提升评估可靠性。TurtleBench包含1,532条用户猜测及标注正确性。我们据此全面评估了当前九个最先进的大语言模型。值得注意的是,OpenAI o1系列模型在该评测中未取得领先结果。我们提出若干假设供后续研究,如“o1的潜在推理可能使用了简单的思维链(CoT)技巧”以及“增加思维链长度虽有推理益处,但也带来噪声代价”。
原文摘要 · Abstract (English)
As the application of Large Language Models (LLMs) expands, the demand for reliable evaluations increases. Existing LLM evaluation benchmarks primarily rely on static datasets, making it challenging to assess model performance in dynamic interactions with users. Moreover, these benchmarks often depend on specific background knowledge, complicating the measurement of a model's logical reasoning capabilities. Other dynamic evaluation methods based on strong models or manual efforts may introduce biases and incur high costs and time demands, hindering large-scale application. To address these issues, we propose TurtleBench. TurtleBench collects real user guesses from our online Turtle Soup Puzzle platform that we developed. This approach allows for the relatively dynamic generation of evaluation datasets, mitigating the risk of model cheating while aligning assessments more closely with genuine user needs for reasoning capabilities, thus enhancing the reliability of evaluations. TurtleBench includes 1,532 user guesses along with the correctness of guesses after annotation. Using this dataset, we thoroughly evaluated nine of the most advanced LLMs available today. Notably, the OpenAI o1 series models did not achieve leading results in these evaluations. We propose several hypotheses for further research, such as "the latent reasoning of o1 utilizes trivial Chain-of-Thought (CoT) techniques" and "increasing CoT length not only provides reasoning benefits but also incurs noise costs."
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。