arXiv:2509.16722cs.CL2025-09EMNLP被引 4

构建多层级因果理解数据集,专攻社交媒体中的隐含因果关系。

A Multi-Level Benchmark for Causal Language Understanding in Social Media Discourse

  • 基于五年Reddit帖子构建多任务标注数据集
  • 涵盖显性/隐性因果识别与因果摘要生成等四类任务
  • 融合专家与大模型标注,支持判别与生成模型评估

理解非正式话语中的因果语言是自然语言处理中的核心挑战,但研究仍不充分。现有数据集主要关注结构化文本中的显性因果,难以支撑对社交媒体用户生成内容中隐含因果表达的检测。我们提出CausalTalk,一个包含2020至2024年五年间关于新冠疫情公共健康议题的Reddit帖子数据集,共10120篇经过标注,覆盖四项因果任务:(1)二分类因果判断,(2)显性与隐性因果区分,(3)因果作用范围抽取,(4)因果主旨生成。标注包括领域专家提供的黄金标准标签及由GPT-4o生成并经人工验证的银标准标签。CausalTalk弥合了细粒度因果检测与基于主旨的推理之间的差距,支持判别式与生成式模型的联合基准测试,并为社交语境下因果推理研究提供丰富资源。

原文摘要 · Abstract (English)

Understanding causal language in informal discourse is a core yet underexplored challenge in NLP. Existing datasets largely focus on explicit causality in structured text, providing limited support for detecting implicit causal expressions, particularly those found in informal, user-generated social media posts. We introduce CausalTalk, a multi-level dataset of five years of Reddit posts (2020-2024) discussing public health related to the COVID-19 pandemic, among which 10120 posts are annotated across four causal tasks: (1) binary causal classification, (2) explicit vs. implicit causality, (3) cause-effect span extraction, and (4) causal gist generation. Annotations comprise both gold-standard labels created by domain experts and silver-standard labels generated by GPT-4o and verified by human annotators. CausalTalk bridges fine-grained causal detection and gist-based reasoning over informal text. It enables benchmarking across both discriminative and generative models, and provides a rich resource for studying causal reasoning in social media contexts.

因果推理社交媒体多任务数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。