构建2.5万条佛兰德语日常叙事数据集,助力低资源语言NLP研究
FLAME: A New Dataset on FLemish Accounts of Momentary Experiences
- 通过体验取样收集佛兰德语个人叙事,构建低资源语言语料库
- 人类评估显示BERTopic在主题一致性与文化相关性上优于其他方法
- 适合研究低资源语言、文化语境下日常语言使用的学者使用
我们提出FLAME(佛兰德语即时体验叙述)语料库,包含近2.5万条比利时荷兰语(佛兰德语)的个人叙事,通过体验取样方式收集,旨在支持对低资源语言的自然语言处理研究。这些日常叙述富含文化相关主题,但其非正式语体和低资源特性使主题提取困难。对比K-Means、LDA和BERTopic三种方法,人工评估表明BERTopic生成的主题最连贯且最具文化共鸣。因此,FLAME为研究低资源语言中日常语言使用提供了新资源。
原文摘要 · Abstract (English)
We introduce FLAME (FLemish Accounts of Momentary Experiences), a corpus of nearly 25,000 personal narratives in Belgian-Dutch (Flemish), collected through experience sampling to support Natural Language Processing (NLP) research on an underrepresented variety. Such everyday narratives are rich in culturally grounded themes, but their informal register and low-resource setting make thematic extraction hard. Comparing K-Means, LDA, and BERTopic, we find that human evaluation favors BERTopic, which produces the most coherent, culturally resonant topics. FLAME, thereby, offers a new resource for studying everyday language use in a low-resource variety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。