arXiv:2601.13253cs.CLcs.LG2026-01

用混合方法低成本构建大规模土耳其语语义数据集

A Hybrid Protocol for Large-Scale Semantic Dataset Generation in Low-Resource Languages: The Turkish Semantic Relations Corpus

  • 结合词向量聚类与大模型自动标注生成语义对
  • 产出84.3万条语义关系,较现有资源提升10倍
  • 适合低资源语言NLP研究者参考使用

我们提出一种混合方法,用于在低资源语言中生成大规模语义关系数据集,以土耳其语语义关系语料库为例进行验证。该方法包含三个阶段:(1) 使用FastText嵌入与凝聚聚类识别语义簇;(2) 采用Gemini 2.5-Flash自动化分类语义关系;(3) 融合人工校验词典源。最终数据集包含843,000条独特的土耳其语语义对,涵盖同义、反义、共下位三种关系类型,规模较现有资源提升10倍,成本仅65美元。通过两个下游任务验证:嵌入模型达到90%的top-1检索准确率,分类模型F1-macro达90%。该可扩展协议有效缓解土耳其语NLP的数据稀缺问题,并具备向其他低资源语言推广的潜力。数据集与模型已公开发布。

原文摘要 · Abstract (English)

We present a hybrid methodology for generating large-scale semantic relationship datasets in low-resource languages, demonstrated through a comprehensive Turkish semantic relations corpus. Our approach integrates three phases: (1) FastText embeddings with Agglomerative Clustering to identify semantic clusters, (2) Gemini 2.5-Flash for automated semantic relationship classification, and (3) integration with curated dictionary sources. The resulting dataset comprises 843,000 unique Turkish semantic pairs across three relationship types (synonyms, antonyms, co-hyponyms) representing a 10x scale increase over existing resources at minimal cost ($65). We validate the dataset through two downstream tasks: an embedding model achieving 90% top-1 retrieval accuracy and a classification model attaining 90% F1-macro. Our scalable protocol addresses critical data scarcity in Turkish NLP and demonstrates applicability to other low-resource languages. We publicly release the dataset and models.

语义关系低资源语言数据集构建土耳其语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。