arXiv:2608.04847cs.CL2026-08

构建首个用户生成的酷儿俚语数据集,评估大模型对边缘群体语言的理解能力。

Do Language Models Know Their Slang? Queer Slang Understanding in User-Generated Content

论文配图:Do Language Models Know Their Slang? Queer Slang Understanding in User-Generated Content
图 1 · 摘自论文原文
  • 人工构建包含118个酷儿术语的语料库与定义
  • 首次测试大模型在不同提示下识别和解释酷儿俚语的能力
  • 为研究敏感身份语言处理提供可复用基准

尽管酷儿俚语具有重要的文化意义并广泛传播,但在自然语言处理研究中仍被严重忽视。为填补这一空白,我们提出了Slang-Q——一个由人工标注的英文用户生成句子数据集,包含118个酷儿术语及其参考定义,并基于新构建的术语分类体系。利用该资源,我们首次探索了语言模型在不同提示条件下理解与定义酷儿俚语的能力。Slang-Q旨在作为研究当前模型如何处理敏感、社群特定语言的基础,评估其对身份表达与语言形式提供准确可靠信息的能力。

原文摘要 · Abstract (English)

Despite its cultural relevance and diffusion, queer slang remains underrepresented in Natural Language Processing research. Towards addressing this gap, we introduce Slang-Q, a manually curated dataset of naturally user-generated English sentences paired with queer slang terms and reference definitions, built upon a newly constructed taxonomy of 118 queer terms. We use this resource to conduct a first exploratory evaluation of language models on their ability to understand and define queer slang under varying prompting conditions. Slang-Q is intended as a basis for studying how current models handle sensitive, community-specific language and whether they can provide accurate and reliable information about such forms of identity and linguistic expression.

酷儿语言大模型评估社会敏感性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。