arXiv:2509.26160cs.CL2025-09被引 3

构建超400万条自然语境通用句数据集,推动通用性语言研究

MGen: Millions of Naturally Occurring Generics in Context

  • 从网页和论文中提取400万条带长上下文的通用句
  • 平均句长超16词,多用于表达对人群的普遍判断
  • 首个大规模多样化通用句数据集,适合语言学与NLP研究

MGen 是一个包含超过400万条自然生成的通用句和量化句的数据集,源自多样文本来源。句子附带长上下文(对应网站和学术论文),涵盖11种不同量化词。分析显示,通用句可长达16词以上,说话者常用于表达对人的普遍性判断。MGen是目前最大、最多样化的自然通用句数据集,为大规模计算研究通用性打开新门。数据集已公开:https://gustavocilleruelo.com/mgen

原文摘要 · Abstract (English)

MGen is a dataset of over 4 million naturally occurring generic and quantified sentences extracted from diverse textual sources. Sentences in the dataset have long context documents, corresponding to websites and academic papers, and cover 11 different quantifiers. We analyze the features of generics sentences in the dataset, with interesting insights: generics can be long sentences (averaging over 16 words) and speakers often use them to express generalisations about people. MGen is the biggest and most diverse dataset of naturally occurring generic sentences, opening the door to large-scale computational research on genericity. It is publicly available at https://gustavocilleruelo.com/mgen

通用句数据集自然语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。