arXiv:2505.15071cs.CL2025-05ACL被引 3

用用户生成内容训练大模型理解中文网络热词。

Can Large Language Models Understand Internet Buzzwords Through User-Generated Content

  • 提出新方法RESS,让大模型像人一样学习热词含义。
  • 构建首个中文热词数据集CHEER,含定义与真实语料。
  • 发现大模型普遍依赖已有知识,推理能力不足。

中文社交媒体的海量用户生成内容(UGC)为研究网络热词提供了可能。本文探讨大语言模型(LLMs)能否基于UGC示例生成准确的热词定义。工作贡献有三:第一,构建了首个中文热词数据集CHEER,每个热词均配有定义和相关UGC;第二,提出新型方法RESS,有效引导LLM的语义理解过程,模拟人类语言学习能力;第三,利用CHEER对多种现成定义生成方法及RESS进行基准测试。结果表明RESS表现更优,但揭示出共性挑战:过度依赖先验知识、推理能力薄弱、难以识别高质量UGC以辅助理解。本研究为基于LLM的定义生成奠定了基础。数据集与代码见https://github.com/SCUNLP/Buzzword。

原文摘要 · Abstract (English)

The massive user-generated content (UGC) available in Chinese social media is giving rise to the possibility of studying internet buzzwords. In this paper, we study if large language models (LLMs) can generate accurate definitions for these buzzwords based on UGC as examples. Our work serves a threefold contribution. First, we introduce CHEER, the first dataset of Chinese internet buzzwords, each annotated with a definition and relevant UGC. Second, we propose a novel method, called RESS, to effectively steer the comprehending process of LLMs to produce more accurate buzzword definitions, mirroring the skills of human language learning. Third, with CHEER, we benchmark the strengths and weaknesses of various off-the-shelf definition generation methods and our RESS. Our benchmark demonstrates the effectiveness of RESS while revealing crucial shared challenges: over-reliance on prior exposure, underdeveloped inferential abilities, and difficulty identifying high-quality UGC to facilitate comprehension. We believe our work lays the groundwork for future advancements in LLM-based definition generation. Our dataset and code are available at https://github.com/SCUNLP/Buzzword.

大模型热词理解用户生成内容

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。