用AI辅助众包构建大规模逻辑谬误新闻评论数据集
CoCoLoFa: A Dataset of News Comments with Common Logical Fallacies Written by LLM-Assisted Crowds
- 让众包工作者在AI助手帮助下撰写含特定谬误的评论
- 数据集含7706条评论,模型在测试集上准确率达86%
- 适合研究虚假信息检测与AI辅助内容生成的人
检测文本中的逻辑谬误有助于用户识别论证缺陷,但自动化检测难度较大。人工标注大规模真实语料中的谬误以构建训练数据成本高昂。本文提出CoCoLoFa,目前最大的逻辑谬误数据集,包含648篇新闻文章的7,706条评论,每条评论均标注了谬误是否存在及类型。我们招募143名众包工作者,在新闻文章背景下撰写体现特定谬误类型(如滑坡谬误)的评论。为应对写作难度,我们在工作界面中集成LLM辅助工具,协助撰写与润色。专家评估显示其文本质量与标注可靠性高。基于BERT的模型在该数据集上微调后,检测与分类的F1值分别达到0.86和0.87,优于当前最先进的LLM。研究表明,结合众包与大模型能更高效地构建复杂语言现象的数据集。
原文摘要 · Abstract (English)
Detecting logical fallacies in texts can help users spot argument flaws, but automating this detection is not easy. Manually annotating fallacies in large-scale, real-world text data to create datasets for developing and validating detection models is costly. This paper introduces CoCoLoFa, the largest known logical fallacy dataset, containing 7,706 comments for 648 news articles, with each comment labeled for fallacy presence and type. We recruited 143 crowd workers to write comments embodying specific fallacy types (e.g., slippery slope) in response to news articles. Recognizing the complexity of this writing task, we built an LLM-powered assistant into the workers' interface to aid in drafting and refining their comments. Experts rated the writing quality and labeling validity of CoCoLoFa as high and reliable. BERT-based models fine-tuned using CoCoLoFa achieved the highest fallacy detection (F1=0.86) and classification (F1=0.87) performance on its test set, outperforming the state-of-the-art LLMs. Our work shows that combining crowdsourcing and LLMs enables us to more effectively construct datasets for complex linguistic phenomena that crowd workers find challenging to produce on their own.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。