对比真实与创作的社交媒体情感数据,发现两者在表达方式上差异明显。
Donate or Create? Comparing Data Collection Strategies for Emotion-labeled Multimodal Social Media Posts
- 用真实投稿和实验创作两种方式收集情感标注的多模态数据
- 创作内容更长、依赖文本、聚焦典型情绪事件,但模型泛化需真实数据验证
- 适合研究数据采集方法影响或评估模型在真实场景表现的学者
准确建模主观现象(如情绪表达)需要标注作者意图的数据。通常通过让参与者捐赠真实世界生成的内容并自行标注,或在研究中创作符合特定标签的内容来获取。后者实施更简便,对参与者隐私风险更低。但尚不清楚研究创作内容与真实内容的差异及其对模型的影响。我们收集了研究创作和真实来源的多模态社交媒体帖子,并从多个维度进行比较,包括模型性能。结果发现:相比真实帖子,研究创作的帖子更长,更依赖文本而非图像传递情绪,且更关注情绪典型事件。愿意捐赠与愿意创作的参与者在人口统计特征上存在差异。研究创作数据有助于训练在真实数据上表现良好的模型,但对模型真实效果的评估仍需依赖真实数据。
原文摘要 · Abstract (English)
Accurate modeling of subjective phenomena such as emotion expression requires data annotated with authors' intentions. Commonly such data is collected by asking study participants to donate and label genuine content produced in the real world, or create content fitting particular labels during the study. Asking participants to create content is often simpler to implement and presents fewer risks to participant privacy than data donation. However, it is unclear if and how study-created content may differ from genuine content, and how differences may impact models. We collect study-created and genuine multimodal social media posts labeled for emotion and compare them on several dimensions, including model performance. We find that compared to genuine posts, study-created posts are longer, rely more on their text and less on their images for emotion expression, and focus more on emotion-prototypical events. The samples of participants willing to donate versus create posts are demographically different. Study-created data is valuable to train models that generalize well to genuine data, but realistic effectiveness estimates require genuine data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。