构建大规模习语与修辞语言数据集,助力模型理解非字面语言。
NLP Datasets for Idiom and Figurative Language Tasks
- 从大语料中提取潜在习语表达,构建多类别数据集
- 创建两个人工标注数据集用于识别习语和修辞含义
- 支持零样本与微调训练,适合语言模型评估与改进
习语和修辞语言构成口语和书面语的重要部分。随着社交媒体的发展,这类非正式语言对普通用户和大语言模型训练者都更加可见。尽管大规模语料看似能解决所有自然语言处理问题,但习语和修辞语言仍难以被大模型准确理解。微调方法虽有效,但更高质量、更大规模的数据集可进一步缩小差距。本文提出一个综合数据集,整合多个现有习语与修辞语言数据集,生成统一习语列表,并从中抽取上下文序列。构建了一个包含大量潜在习语表达的大规模数据集,以及两个经人工标注的明确习语与修辞语言数据集,用于评估预训练模型在习语识别任务中的表现。数据集经过后处理以兼容不同模型训练,应用于槽位标注与序列标注任务,验证了其有效性。
原文摘要 · Abstract (English)
Idiomatic and figurative language form a large portion of colloquial speech and writing. With social media, this informal language has become more easily observable to people and trainers of large language models (LLMs) alike. While the advantage of large corpora seems like the solution to all machine learning and Natural Language Processing (NLP) problems, idioms and figurative language continue to elude LLMs. Finetuning approaches are proving to be optimal, but better and larger datasets can help narrow this gap even further. The datasets presented in this paper provide one answer, while offering a diverse set of categories on which to build new models and develop new approaches. A selection of recent idiom and figurative language datasets were used to acquire a combined idiom list, which was used to retrieve context sequences from a large corpus. One large-scale dataset of potential idiomatic and figurative language expressions and two additional human-annotated datasets of definite idiomatic and figurative language expressions were created to evaluate the baseline ability of pre-trained language models in handling figurative meaning through idiom recognition (detection) tasks. The resulting datasets were post-processed for model agnostic training compatibility, utilized in training, and evaluated on slot labeling and sequence tagging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。