arXiv:2511.18622cs.CLcs.AI2025-11

用AI一周生成百万级词典,涵盖定义、用法、词源等

OpenGloss: A Synthetic Encyclopedic Dictionary and Semantic Knowledge Graph

  • 通过多智能体生成流程,自动构建词义与语义关系
  • 含53.7万词义、910万语义边、6000万词百科内容
  • 适合语言学习、NLP研究者快速获取高质量词汇资源

我们提出OpenGloss,一个合成的英语百科词典与语义知识图谱,整合词义定义、百科背景、词源历史及语义关系。该资源包含53.7万个词义(对应15万个词素),与WordNet 3.1和Open English WordNet相当,但词义定义数量超四倍。涵盖910万条语义边、100万条用例、300万组搭配词,以及6000万词的百科内容。通过基于模式验证的LLM多智能体生成流水线与自动化质量控制,在不到一周时间内以不足1000美元完成全部生成。这证明结构化生成可在人力无法企及的成本与时间下构建综合性词汇资源,支持基础模型迭代。其整合定义、例句、搭配、词源与百科内容,填补教学应用空白,既可用于词汇学习,也适用于自然语言处理任务。作为合成数据,OpenGloss反映当前基础模型的能力与局限。数据集已公开于Hugging Face,采用CC-BY 4.0许可,供研究者与教育工作者使用。

原文摘要 · Abstract (English)

We present OpenGloss, a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. OpenGloss contains 537K senses across 150K lexemes, on par with WordNet 3.1 and Open English WordNet, while providing more than four times as many sense definitions. These lexemes include 9.1M semantic edges, 1M usage examples, 3M collocations, and 60M words of encyclopedic content. Generated through a multi-agent procedural generation pipeline with schema-validated LLM outputs and automated quality assurance, the entire resource was produced in under one week for under $1,000. This demonstrates that structured generation can create comprehensive lexical resources at cost and time scales impractical for manual curation, enabling rapid iteration as foundation models improve. The resource addresses gaps in pedagogical applications by providing integrated content -- definitions, examples, collocations, encyclopedias, etymology -- that supports both vocabulary learning and natural language processing tasks. As a synthetically generated resource, OpenGloss reflects both the capabilities and limitations of current foundation models. The dataset is publicly available on Hugging Face under CC-BY 4.0, enabling researchers and educators to build upon and adapt this resource.

词典生成知识图谱合成数据语言学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。