测试大模型在韩语文化背景下的多步推理能力,发现性能存在突然跃升现象。
Multi-Step Reasoning in Korean and the Emergent Mirage
- 用模板生成韩语文化相关多步推理题,考察模型跨步骤理解能力。
- 训练量低于2×10²⁵ FLOPs时几乎零分,超过阈值后性能急剧提升。
- 性能跃升可能源于错误累积而非真正新能力,适合研究模型可靠性者关注。
我们提出HRMCR(HAE-RAE多步常识推理)基准,用于评估大语言模型在特定文化语境——韩国文化中的多步推理能力。题目通过模板与算法自动生成,要求模型将韩式文化知识融入连续推理步骤中。实验显示,训练量少于2×10²⁵ FLOPs的模型几乎无法解答任何问题,表现接近零分;超过该阈值后,性能迅速上升。当前最先进模型(如O1)得分仍不足50%,凸显任务难度。值得注意的是,逐步分析表明,观察到的涌现行为可能源于多步推理中错误的累积,而非真正的新能力。我们已公开发布该基准,并承诺定期更新数据集以防止污染。
原文摘要 · Abstract (English)
We introduce HRMCR (HAE-RAE Multi-Step Commonsense Reasoning), a benchmark designed to evaluate large language models' ability to perform multi-step reasoning in culturally specific contexts, focusing on Korean. The questions are automatically generated via templates and algorithms, requiring LLMs to integrate Korean cultural knowledge into sequential reasoning steps. Consistent with prior observations on emergent abilities, our experiments reveal that models trained on fewer than \(2 \cdot 10^{25}\) training FLOPs struggle to solve any questions, showing near-zero performance. Beyond this threshold, performance improves sharply. State-of-the-art models (e.g., O1) still score under 50\%, underscoring the difficulty of our tasks. Notably, stepwise analysis suggests the observed emergent behavior may stem from compounding errors across multiple steps rather than reflecting a genuinely new capability. We publicly release the benchmark and commit to regularly updating the dataset to prevent contamination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。