用大模型生成英捷语语料,可与真实语料直接对比研究。
AI Brown and AI Koditex: LLM-Generated Corpora Comparable to Traditional Corpora of English and Czech Texts
- 用多款大模型生成英捷语语料,覆盖多种文体和主题。
- 英语文本平均每模型86.4万词元,总计2700万;捷语文本共2150万词元。
- 语料已标注词性、句法等,适合语言学与NLP研究者使用。
本文介绍了两个由大语言模型生成的英语和捷克语语料库。其目的是创建资源,用于从语言学角度比较人类写作与大模型生成文本。语料库注重多文体、广主题、多作者与多文本类型,并确保与现有真人语料库具有可比性。生成语料库复现了参考语料库:英语的BE21(保罗·贝克尔的现代版布朗语料库)与捷克语的Koditex语料库,均遵循布朗语料库传统。生成使用OpenAI、Anthropic、Alphabet、Meta和DeepSeek的模型,涵盖GPT-3(davinci-002)至GPT-4.5。所有数据按统一依存标注标准(Universal Dependencies)进行分词、词形还原及语法句法标注。英语文本平均每模型86.4万词元,总计2700万词元;捷语文本平均每模型76.8万词元,总计2150万词元。语料库免费下载(CC BY 4.0),标注数据采用CC BY-NC-SA 4.0许可,也可通过捷克国家语料库搜索接口访问。
原文摘要 · Abstract (English)
This article presents two corpora of English and Czech texts generated with large language models (LLMs). The motivation is to create a resource for comparing human-written texts with LLM-generated text linguistically. Emphasis was placed on ensuring these resources are multi-genre and rich in terms of topics, authors, and text types, while maintaining comparability with existing human-created corpora. These generated corpora replicate reference human corpora: BE21 by Paul Baker, which is a modern version of the original Brown Corpus, and Koditex corpus that also follows the Brown Corpus tradition but in Czech. The new corpora were generated using models from OpenAI, Anthropic, Alphabet, Meta, and DeepSeek, ranging from GPT-3 (davinci-002) to GPT-4.5, and are tagged according to the Universal Dependencies standard (i.e., they are tokenized, lemmatized, and morphologically and syntactically annotated). The subcorpus size varies according to the model used (the English part contains on average 864k tokens per model, 27M tokens altogether, the Czech partcontains on average 768k tokens per model, 21.5M tokens altogether). The corpora are freely available for download under the CC BY 4.0 license (the annotated data are under CC BY-NC-SA 4.0 licence) and are also accessible through the search interface of the Czech National Corpus.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。