评测大模型对中文新词的理解与还原能力,发现多数模型表现不佳。
CNeo-Bench: Diagnosing Large Language Models on Chinese Neologisms

- 构建4759个中文新词的基准测试,按语言机制分类
- 多数模型定义生成准确率低于40%,存在识别与还原的差距
- 提示工程可缓解部分难题,但根本挑战仍存
中文新词运用多种独特语言机制,如谐音替换(如‘886’代指‘拜拜’)和字形拆解,这些在其他语言中罕见。本文提出 CNeo-Bench,一个包含 4,759 个中文新词及其参考定义的基准,按语言机制分为五大类九小类。该基准配套双层评估框架,区分模型是否能描述新词与是否能操作其底层机制。对 18 个 LLM 的评估显示,中文新词仍是开放挑战:多数模型定义生成准确率低于 40%;在多个子类别中出现系统性“识别-还原”差距——模型能正确描述新词,但在源形式还原任务中仅输出语义等价表达,而非原始形式。对 1,058 个难题样本的少样本分析表明,上下文示例可解决许多困难案例,但仍残留显著错误,说明挑战远非仅靠提示工程可克服。
原文摘要 · Abstract (English)
Chinese neologisms exploit diverse and unique linguistic mechanisms, such as phonetic substitution (e.g., 886 for ``bye-bye'') and visual character decomposition that are rare in other languages. We introduce CNeo-Bench, a benchmark of 4,759 such neologisms with reference definitions, organized into five top-level categories and nine subcategories by the linguistic mechanism behind each expression. CNeo-Bench is paired with a two-tier evaluation framework that separates whether a model can describe a neologism from whether it can operate on its underlying mechanism. Evaluating 18 LLMs, we find that Chinese neologisms remain an open challenge; most models fall below 40\% on definition generation, and on several subcategories a systematic recognition-manipulation gap emerges: models describe neologisms correctly but, in source-form restoration tasks, substitute a semantic equivalent (paraphrase) for the source form rather than producing the source form itself. A few-shot analysis on 1,058 hard items shows that in-context examples can solve many difficult cases, but leave a noticeable portion of errors remaining, indicating challenges beyond prompting alone can address.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。