构建5.5万条梦境报告库,助力清醒梦研究。
A large corpus of lucid and non-lucid dream reports
- 从在线论坛收集5000人十年的匿名梦境记录
- 包含1万条清醒梦、2.5万条普通梦及2千条噩梦标注
- 语言特征验证显示清醒梦报告符合已知现象学特征
所有类型的梦境仍属未知。特别是具有梦境中自我意识的清醒梦,因其稀有性和难以主动诱发,导致高质量梦境报告数量有限。这一困境使得清醒梦的主观体验特征不明确,其诸多潜在应用未能充分实现。本文构建了一个包含5.5万条梦境报告的大型语料库,来自5000名参与者,数据源自一个公开的在线论坛,用户可自主标记梦境为清醒梦、非清醒梦或噩梦。其中包含1万条清醒梦标签、2.5万条非清醒梦标签和2千条噩梦标签。通过描述性统计与可视化分析,并进行结构效度检验,发现清醒梦标签报告的语言模式与已知清醒梦特征一致。整个语料库对梦境科学研究具有广泛价值,而标注子集尤其适用于新发现的清醒梦研究。
原文摘要 · Abstract (English)
All varieties of dreaming remain a mystery. Lucid dreams in particular, or those characterized by awareness of the dream, are notoriously difficult to study. Their scarce prevalence and resistance to deliberate induction make it difficult to obtain a sizeable corpus of lucid dream reports. The consequent lack of clarity around lucid dream phenomenology has left the many purported applications of lucidity under-realized. Here, a large corpus of 55k dream reports from 5k contributors is curated, described, and validated for future research. Ten years of publicly available dream reports were scraped from an online forum where users share anonymous dream journals. Importantly, users optionally categorize their dream as lucid, non-lucid, or a nightmare, offering a user-provided labeling system that includes 10k lucid and 25k non-lucid, and 2k nightmare labels. After characterizing the corpus with descriptive statistics and visualizations, construct validation shows that language patterns in lucid-labeled reports are consistent with known characteristics of lucid dreams. While the entire corpus has broad value for dream science, the labeled subset is particularly powerful for new discoveries in lucid dream studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。