arXiv:2507.03488cs.CL2025-07

构建生命科学领域虚假信息数据集,助力识别误导性文本。

Four Shades of Life Sciences: A Dataset for Disinformation Detection in the Life Sciences

  • 基于语言与修辞特征,区分虚假与真实生命科学文本。
  • 创建含2603篇文本的FSoLS数据集,覆盖14个主题与4类出版物。
  • 开源可复现,适合研究虚假信息检测与健康传播的学者使用。

虚假信息传播者常通过吸引关注或激发情绪来获取影响力或收益,形成独特的修辞模式,可被机器学习模型利用。本文探讨语言与修辞特征作为区分误导性文本与其他生命科学文本的代理指标,结合大语言模型与传统机器学习分类器进行分析。针对现有数据集多聚焦事实核查而忽略语境的局限,我们提出新的标注语料库Four Shades of Life Sciences (FSoLS),包含2,603篇来自17个不同来源、涵盖14个生命科学主题的文本,并划分为四类生命科学出版物。相关源代码已开源,可在GitHub上获取:https://github.com/EvaSeidlmayer/FourShadesofLifeSciences。

原文摘要 · Abstract (English)

Disseminators of disinformation often seek to attract attention or evoke emotions - typically to gain influence or generate revenue - resulting in distinctive rhetorical patterns that can be exploited by machine learning models. In this study, we explore linguistic and rhetorical features as proxies for distinguishing disinformative texts from other health and life-science text genres, applying both large language models and classical machine learning classifiers. Given the limitations of existing datasets, which mainly focus on fact checking misinformation, we introduce Four Shades of Life Sciences (FSoLS): a novel, labeled corpus of 2,603 texts on 14 life-science topics, retrieved from 17 diverse sources and classified into four categories of life science publications. The source code for replicating, and updating the dataset is available on GitHub: https://github.com/EvaSeidlmayer/FourShadesofLifeSciences

虚假信息生命科学文本分类数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。