arXiv:2601.19932cs.CLcs.HC2026-01ACL

构建中文网络评论中隐晦表达的分类体系与数据集,揭示大模型理解难题。

"Newspaper Eat" Means "Not Tasty": A Taxonomy and Benchmark for Coded Language in Real-World Chinese Online Reviews

  • 提出7类编码策略分类法,涵盖音近、形似等隐晦表达方式。
  • 在7744条真实评论上验证,强模型仍频繁误判编码语义。
  • 适合关注中文网络语义理解、社会语言学的研究者使用。

隐晦表达是人类交流的重要组成部分,指用户故意通过表面文字与实际含义的差异来传递信息,需解码才能理解。当前语言模型对这类表达处理不佳,进展受限于真实世界数据集和清晰分类体系的缺乏。本文提出CodedLang数据集,包含7744条来自Google Maps的中文评论,其中900条带有词元级标注。我们构建了一个七类编码策略分类体系,涵盖音近、形近及跨语言替换等常见手法。在编码识别、分类及评分预测任务上对主流语言模型进行基准测试,结果显示即使强模型也常无法识别或正确理解编码内容。由于多数编码依赖发音特征,我们进一步开展了编码与解码形式的语音分析。代码与数据集已公开。研究凸显了隐晦表达作为真实世界NLP系统重要且未被充分探索的挑战。

原文摘要 · Abstract (English)

Coded language is an important part of human communication. It refers to cases where users intentionally encode meaning so that the surface text differs from the intended meaning and must be decoded to be understood. Current language models handle coded language poorly. Progress has been limited by the lack of real-world datasets and clear taxonomies. This paper introduces CodedLang, a dataset of 7,744 Chinese Google Maps reviews, including 900 reviews with span-level annotations of coded language. We developed a seven-class taxonomy that captures common encoding strategies, including phonetic, orthographic, and cross-lingual substitutions. We benchmarked language models on coded language detection, classification, and review rating prediction. Results show that even strong models can fail to identify or understand coded language. Because many coded expressions rely on pronunciation-based strategies, we further conducted a phonetic analysis of coded and decoded forms. Our code and dataset are publicly available. Together, our results highlight coded language as an important and underexplored challenge for real-world NLP systems.

隐晦表达中文NLP数据集语义理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。