arXiv:2604.15841cs.CL2026-04ACL被引 2

测试大模型对中文抽象语言的掌握能力,发现现有模型表现有限。

Exploring the Capability Boundaries of LLMs in Mastering of Chinese Chouxiang Language

论文配图:Exploring the Capability Boundaries of LLMs in Mastering of Chinese Chouxiang Language
图 1 · 摘自论文原文
  • 构建专用基准Mouse,评估大模型在六类任务中的表现。
  • 当前顶尖模型在抽象语言任务上普遍表现不佳,仅在语义理解中较好。
  • 揭示模型翻译偏差原因,适合关注网络亚文化与多文化NLP的研究者。

尽管大语言模型(LLMs)在通用语言任务中取得显著进展,但其在中文网络亚文化语言——抽象话(Chouxiang Language)上的表现仍鲜有研究。本文提出Mouse,一个专门用于评估LLMs在六类涉及抽象话的自然语言处理任务中的基准。实验表明,当前最先进的(SOTA)LLMs在多项任务中存在明显局限,但在上下文语义理解任务中表现尚可。此外,我们探讨了SOTA模型在抽象话上表现不佳的原因,检验了翻译任务中采用的LLM作为裁判方法是否符合人类判断,并分析影响抽象话翻译的关键因素。本研究旨在推动自然语言处理领域对多元文化融合及动态演变网络语言的进一步探索。代码与数据已公开。

原文摘要 · Abstract (English)

While large language models (LLMs) have achieved remarkable success in general language tasks, their performance on Chouxiang Language, a representative subcultural language in the Chinese internet context, remains largely unexplored. In this paper, we introduce Mouse, a specialized benchmark designed to evaluate the capabilities of LLMs on NLP tasks involving Chouxiang Language across six tasks. Experimental results show that, current state-of-the-art (SOTA) LLMs exhibit clear limitations on multiple tasks, while performing well on tasks that involve contextual semantic understanding. In addition, we further discuss the reasons behind the generally low performance of SOTA LLMs on Chouxiang Language, examine whether the LLM-as-a-judge approach adopted for translation tasks aligns with human judgments and values, and analyze the key factors that influence Chouxiang translation. Our study aims to promote further research in the NLP community on multicultural integration and the dynamics of evolving internet languages. Our code and data are publicly available.

大模型评测中文亚文化抽象话NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。