arXiv:2602.14406cs.CLcs.AI2026-02

构建了真相社交平台对话数据集,用于研究网络观点争辩与立场识别。

TruthStance: An Annotated Dataset of Conversations on Truth Social

  • 构建包含24,378篇帖子和52万条评论的真言社交对话数据集
  • 标注1500条实例,评估大模型提示策略在立场检测上的表现
  • 提供大模型生成的立场与论点标签,支持多维度分析

论点挖掘与立场识别是理解在线话语中观点形成与争议的核心。然而,现有公开资源多集中于推特、Reddit等主流平台,对替代性技术平台的对话结构研究不足。本文引入TruthStance,一个涵盖2023-2025年真言社交(Truth Social)对话线程的大规模数据集,包含24,378篇帖子和523,360条评论,并保留回复树结构。我们提供人工标注基准数据集1,500个实例,涵盖论点挖掘与基于主张的立场检测,含标注者间一致性指标,并用于评估大语言模型(LLM)提示策略。采用最优配置后,释放额外的LLM生成标签:24,352篇帖子(论点存在性)和107,873条评论(对父节点的立场),支持跨深度、主题与用户群体的立场与论证模式分析。所有代码与数据均公开可用。

原文摘要 · Abstract (English)

Argument mining and stance detection are central to understanding how opinions are formed and contested in online discourse. However, most publicly available resources focus on mainstream platforms such as Twitter and Reddit, leaving conversational structure on alt-tech platforms comparatively under-studied. We introduce TruthStance, a large-scale dataset of Truth Social conversation threads spanning 2023-2025, consisting of 24,378 posts and 523,360 comments with reply-tree structure preserved. We provide a human-annotated benchmark of 1,500 instances across argument mining and claim-based stance detection, including inter-annotator agreement, and use it to evaluate large language model (LLM) prompting strategies. Using the best-performing configuration, we release additional LLM-generated labels for 24,352 posts (argument presence) and 107,873 comments (stance to parent), enabling analysis of stance and argumentation patterns across depth, topics, and users. All code and data are released publicly.

立场识别对话数据大模型论点挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。