构建多语言逻辑谜题数据集,评估大模型跨语言推理能力。
MultiZebraLogic: A Multilingual Logical Reasoning Benchmark
- 设计九语种逻辑谜题,含干扰信息提升难度。
- 2×3与4×5规模谜题分别适配非推理与推理模型。
- 发现语言、文化主题和线索类型对模型表现无显著影响。
我们构建了九种语言的高质量逻辑推理评估数据集,由母语者人工校验。数据集采用名为“zebra puzzles”的逻辑谜题,并分析了多种调节难度的方法,包括谜题规模(对象数与线索数)以及新增的干扰线索(仅含无关信息)。实验表明,干扰线索确实显著增加了模型解题难度;2×3规模的谜题对GPT-4o mini(非推理模型)已具挑战性,而4×5规模对o3-mini(推理模型)亦足够困难。我们进一步分析了语言、文化主题敏感性及线索类型对模型性能的影响,使用英语与丹麦语进行测试,结果表明在OpenAI的GPT-4o mini与o3-mini模型上,三者均无显著差异。研究公开了九种语言下2×3与4×5规模的完整数据集,以及生成谜题的代码,支持扩展至更多语言。
原文摘要 · Abstract (English)
We create high-quality datasets for LLM evaluation of logical reasoning skills across nine different languages, which have been manually checked by fluent speakers. The datasets consist of so-called zebra puzzles, and we analyse different ways of tuning the difficulty of the puzzles to fit modern LLMs. This includes the size of the puzzle (number of objects and number of clues), as well as a novel addition of red herring clues containing only irrelevant information. We show that presence of red herrings indeed makes the puzzles significantly harder for the models, and we find puzzle sizes 2x3 and 4x5 are sufficiently challenging for GPT-4o mini (a non-reasoning model) and o3-mini (a reasoning model), respectively. We analyse whether LLM performance of these are sensitive to the language, the cultural sensitivity of the puzzle theme, and the choice of clue types. These analyses are conducted with English and Danish, where we show that there is no significant difference for either of these three aspects, at least for the OpenAI models GPT-4o mini and o3-mini, chosen as representative non-reasoning and reasoning models, respectively. We publish the datasets for each of the nine languages for the identified sizes 2x3 and 4x5. We also publish the code used to generate the puzzles, which can be used to extend the benchmark into more languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。