arXiv:2606.17905cs.CL2026-06

测试中文表达对逻辑推理模型的影响,发现中英文表现有差距

ChLogic: Evaluating Robustness of Logical Reasoning in Chinese Expressions

论文配图:ChLogic: Evaluating Robustness of Logical Reasoning in Chinese Expressions
图 1 · 摘自论文原文
  • 用中英对照的逻辑题集,检验模型跨语言推理能力
  • 中文表达复杂时,部分模型性能反而下降,翻译未必有效
  • 适合研究多语言模型鲁棒性或中文NLP的学者参考

大语言模型在标准逻辑推理基准上表现日益优异,但其在非英语环境下的鲁棒性尚不明确。我们提出ChLogic,一个中英对齐的基准,用于测试相同潜在逻辑结构在不同中文表面形式下的推理表现是否保持一致。该基准基于形式化逻辑模板构建,包含三个数据集:(i) 通用对齐集,来自九类模板的60个一般命题;(ii) 困难对齐集,来自40个难题;(iii) 中文独有集,涵盖15种语言特异性现象。每个对齐项包含一个英文参考表达和五个中文实现。在Qwen3、Ministral和GLM模型上的实验显示,存在持续的中英文性能差距。从标准中文回译成英文可提升通用对齐集的表现,但在困难对齐集上效果参差,其中Qwen3-32B和GLM-5.1在翻译后表现更差。结果表明,中文表面实现、翻译误差与模型特性共同影响多语言逻辑推理。总体而言,ChLogic为多语言推理的鲁棒性提供了一个有效的压力测试。

原文摘要 · Abstract (English)

Large language models perform increasingly well on standardized logical reasoning benchmarks, but whether this ability remains robust beyond English is unclear. We introduce ChLogic, an English--Chinese aligned benchmark that tests whether models preserve logical reasoning performance when the same latent logical structure is expressed in English and diverse Chinese surface realizations. Built from formal logical templates, the benchmark contains three data sets: (i) the General aligned set, derived from 60 General Propositions across nine template families; (ii) the Difficult aligned set, derived from 40 Difficult Problems; and (iii) the Chinese-only set, covering 15 language-specific phenomenon types. Each aligned item pairs one English reference expression with five Chinese realizations. Experiments on Qwen3, Ministral, and GLM models reveal a persistent English--Chinese performance gap. Back-translation from standard Chinese into English often improves performance on the General aligned set, but produces mixed effects on the Difficult aligned set, where Qwen3-32B and GLM-5.1 perform worse after translation. These results indicate that Chinese surface realization, translation artifacts, and model-specific behavior jointly affect multilingual logical reasoning. Overall, ChLogic provides a useful stress test for the robustness of multilingual reasoning.

逻辑推理多语言中文NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。