构建了3.96万条生物伦理争议的社交媒体立场数据集。
A Context-Aware Dataset for Stance Detection in Bioethical Controversies on Reddit

- 基于Reddit对话构建上下文感知的标注数据集
- 涵盖6个争议主题,标注一致率达0.82(α值)
- 适合研究生物伦理话语与论点挖掘的学者
生物伦理争议在社交媒体上日益增多,但现有立场检测研究缺乏大规模、领域特定的数据资源。本文提出BioStance,一个包含39,600条来自Reddit生物伦理讨论的帖子-评论配对标注数据集。该数据集覆盖三大生物伦理争议维度:基本价值冲突、个人自由与集体责任权衡、技术不确定性,涉及六个争议主题。每条样本保留层级化对话上下文,并由三位独立标注者采用三分类立场标签(支持、反对、无立场)进行标注。标注一致性经计算获得均值Krippendorff's α为0.82,表明标注具有较高可靠性。通过融合主题多样性、对话结构与高质量人工标注,BioStance可支撑上下文感知的立场检测、论点挖掘及生物伦理话语的计算分析研究。
原文摘要 · Abstract (English)
Bioethical debates increasingly unfold on social media, yet stance detection research lacks large-scale, domain-specific resources for modeling such context-dependent discourse. We present BioStance, a context-aware dataset of 39,600 annotated Post-Comment pairs from Reddit bioethical discussions. BioStance covers six controversial targets across three dimensions of bioethical controversy: fundamental value conflicts, individual liberty versus collective responsibility, and technological uncertainty. Each instance preserves hierarchical conversational context and is labeled by three independent annotators using a three-class stance scheme: Favor, Against, and None. The annotations achieve a mean Krippendorff's $α$ of 0.82, indicating substantial reliability. By combining thematic diversity, conversational structure, and high-quality human annotation, BioStance supports research on context-aware stance detection, argument mining, and computational analysis of bioethical discourse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。