构建涵盖五类情感特征的大规模文本数据集,助力跨学科情感研究。
Affect, Body, Cognition, Demographics, and Emotion: The ABCDE of Text Features for Computational Affective Science
- 收集4亿+条来自社交媒体等来源的文本,标注情感、身体、认知等五类特征。
- 提供跨领域可用的标准化数据,支持情绪、行为与社会现象分析。
- 适合心理学、社会学、计算语言学等非计算机背景研究者使用。
计算情感科学与计算社会科学广泛研究人类情绪、行为与健康等相关问题。这类研究通常依赖语言数据,并需先进行标注,如情绪词汇使用或说话人年龄等信息。尽管已有多种资源和算法支持标注,但发现、获取和使用这些工具对计算机科学以外的实践者仍构成显著障碍。本文提出ABCDE数据集(情感、身体、认知、人口统计、情绪),包含超过4亿条来自社交媒体、博客、书籍及人工智能生成内容的文本语料,均标注了适用于计算情感与社会科学研究的多样化特征。该数据集可促进跨学科研究,涵盖情感科学、认知科学、数字人文、社会学、政治学及计算语言学等多个领域。
原文摘要 · Abstract (English)
Work in Computational Affective Science and Computational Social Science explores a wide variety of research questions about people, emotions, behavior, and health. Such work often relies on language data that is first labeled with relevant information, such as the use of emotion words or the age of the speaker. Although many resources and algorithms exist to enable this type of labeling, discovering, accessing, and using them remains a substantial impediment, particularly for practitioners outside of computer science. Here, we present the ABCDE dataset (Affect, Body, Cognition, Demographics, and Emotion), a large-scale collection of over 400 million text utterances drawn from social media, blogs, books, and AI-generated sources. The dataset is annotated with a wide range of features relevant to computational affective and social science. ABCDE facilitates interdisciplinary research across numerous fields, including affective science, cognitive science, the digital humanities, sociology, political science, and computational linguistics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。