对比人与大模型生成俚语,发现模型虽有创意但理解偏差大。
How do Language Models Generate Slang: A Systematic Comparison between Human and Machine-Generated Slang Usages
- 通过三维度评估人类与大模型生成的俚语差异。
- 大模型生成俚语在创意上接近人类,但结构认知存在明显偏差。
- 适合关注大模型语言理解局限的研究者阅读。
俚语是常见的非正式语言形式,对自然语言处理系统构成严峻挑战。尽管大型语言模型(LLMs)的进展使该问题更具可解性,但其在俚语检测与解释等中介任务中的泛化能力与可靠性,取决于模型是否掌握与人类一致的俚语结构知识。为此,本文系统比较了人类与机器生成的俚语用法。评估框架聚焦三个核心方面:1)反映机器对俚语认知系统性偏见的用法特征;2)俚语中体现的词汇创造与词复用所反映的创造力;3)作为模型蒸馏金标准示例时的语义信息量。基于在线俚语词典(OSD)中的人类标注俚语,与GPT-4o和Llama-3生成的俚语进行对比,结果表明大模型在俚语创造性方面已掌握相当知识,但其认知结构与人类显著不一致,难以支持需要外推的任务(如语言学分析)。
原文摘要 · Abstract (English)
Slang is a commonly used type of informal language that poses a daunting challenge to NLP systems. Recent advances in large language models (LLMs), however, have made the problem more approachable. While LLM agents are becoming more widely applied to intermediary tasks such as slang detection and slang interpretation, their generalizability and reliability are heavily dependent on whether these models have captured structural knowledge about slang that align well with human attested slang usages. To answer this question, we contribute a systematic comparison between human and machine-generated slang usages. Our evaluative framework focuses on three core aspects: 1) Characteristics of the usages that reflect systematic biases in how machines perceive slang, 2) Creativity reflected by both lexical coinages and word reuses employed by the slang usages, and 3) Informativeness of the slang usages when used as gold-standard examples for model distillation. By comparing human-attested slang usages from the Online Slang Dictionary (OSD) and slang generated by GPT-4o and Llama-3, we find significant biases in how LLMs perceive slang. Our results suggest that while LLMs have captured significant knowledge about the creative aspects of slang, such knowledge does not align with humans sufficiently to enable LLMs for extrapolative tasks such as linguistic analyses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。