中文日文完美体推理难,大模型表现不佳。
LLMs Struggle with NLI for Perfect Aspect: A Cross-Linguistic Study in Chinese and Japanese
- 构建跨语言完美体NLI数据集,每种语言1350对
- 先进大模型在时态和参照时间变化上错误率高
- 适合研究时态语义与多语言NLP的学者
与英语使用不同形式(如 had、has、will have)标记不同时态的完成体不同,中文和日语在完成体中缺乏独立的时态语法形式,这增加了自然语言推理(NLI)的难度。聚焦于这两种语言的完成体,我们构建了一个基于语言学动机的模板化NLI数据集(每种语言1,350对)。实验表明,即使是最先进的大语言模型在时间推理方面仍表现不佳,尤其在检测细微的时态和参照时间变化时。这些发现揭示了模型的局限性,并强调了在时间语义研究中进行跨语言评估的重要性。我们的数据集已开源:https://github.com/Lujie2001/CrossNLI。
原文摘要 · Abstract (English)
Unlike English, which uses distinct forms (e.g., had, has, will have) to mark the perfect aspect across tenses, Chinese and Japanese lack separate grammatical forms for tense within the perfect aspect, which complicates Natural Language Inference (NLI). Focusing on the perfect aspect in these languages, we construct a linguistically motivated, template-based NLI dataset (1,350 pairs per language). Experiments reveal that even advanced LLMs struggle with temporal inference, particularly in detecting subtle tense and reference-time shifts. These findings highlight model limitations and underscore the need for cross-linguistic evaluation in temporal semantics. Our dataset is available at https://github.com/Lujie2001/CrossNLI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。