测试大模型对中文零代词的理解能力,发现表现普遍不佳。
How Much Do LLMs Know About Chinese Zero Pronouns?

- 设计五类语言学任务评估大模型处理中文零代词的能力
- 即使是顶尖模型,零代词翻译正确率也低于50%
- 上游识别和指代性判断是模型最薄弱环节,适合语言理解研究者参考
零代词(ZPs)是汉语等省略代词语言中的普遍现象,长期给自然语言处理系统带来挑战。尽管大语言模型在多项中文任务中表现良好,但其对零代词的处理能力仍不明确。本文通过一系列语言学驱动的任务——包括识别、指代性分类、指代类型分类、消解与翻译——系统评估了多种大模型对中文零代词的处理能力。实验结果表明,当前大模型在中文零代词处理上仍面临巨大挑战,尤其在上游任务如识别和指代性分类中表现较差。下游任务如零代词翻译性能也普遍偏低:即使是最先进的推理型大模型,正确翻译的零代词也少于一半。
原文摘要 · Abstract (English)
Zero Pronouns (ZPs) are a pervasive linguistic phenomenon in pro-drop languages such as Chinese and have long posed a challenge for natural language processing systems. Although Large Language Models (LLMs) perform well on many Chinese language tasks, their ability to process ZPs remains poorly understood. We conduct a systematic investigation of LLMs' handling of Chinese ZPs through a sequence of linguistically motivated tasks, including identification, referentiality classification, referential type classification, resolution, and translation. A diverse set of LLMs is evaluated across all tasks. Our results show that Chinese ZPs remain highly challenging for current LLMs, particularly for upstream tasks such as identification and referentiality classification. Performance on downstream tasks, such as ZP translation, is also consistently low: even state-of-the-art reasoning-oriented LLMs correctly translate fewer than half of Chinese ZPs into English.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。