300亿词的意大利语论坛数据集,助力语言模型与社会语言学研究。
Testimole-Conversational: A 30-Billion-Word Italian Discussion Board Corpus (1996-2024) for Language Modeling and Sociolinguistic Research
- 构建了1996-2024年意大利语在线论坛语料库,超300亿词
- 覆盖长期语言演变与网络互动模式,支持多维度分析
- 适合语言模型训练、社会语言学及数字沟通研究者使用
我们提出「Testimole-conversational」,一个大规模意大利语讨论板语料库,包含超过300亿词(1996–2024年),适用于意大利语大语言模型预训练。该语料库记录了丰富的计算机中介通信形式,涵盖非正式书面语、话语动态与线上社会互动,具有长期跨度。不仅可用于自然语言处理中的语言建模、领域适配与对话分析,也支持对数字通信中语言变异与社会现象的研究。该资源将免费向学术界开放。
原文摘要 · Abstract (English)
We present "Testimole-conversational" a massive collection of discussion boards messages in the Italian language. The large size of the corpus, more than 30B word-tokens (1996-2024), renders it an ideal dataset for native Italian Large Language Models'pre-training. Furthermore, discussion boards' messages are a relevant resource for linguistic as well as sociological analysis. The corpus captures a rich variety of computer-mediated communication, offering insights into informal written Italian, discourse dynamics, and online social interaction in wide time span. Beyond its relevance for NLP applications such as language modelling, domain adaptation, and conversational analysis, it also support investigations of language variation and social phenomena in digital communication. The resource will be made freely available to the research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。