构建大规模科学事实验证数据集,助力可信科研信息核查
SciClaimHunt: A Large Dataset for Evidence-based Scientific Claim Verification
- 从科研论文中构建双规模数据集,覆盖复杂科学论断
- 提出适配科学论证的基线模型,验证数据集可靠性
- 适合科研人员与AI模型训练者用于可信度评估
科学事实验证比政治或新闻类声明验证更具挑战性。科学论断的受众从研究人员到普通用户差异极大,且证据常涉及专业术语与领域知识,需专用模型支持。尽管研究界兴趣浓厚,仍缺乏大规模可用来训练和评估的科学事实验证数据集。为此,我们从科研论文中构建了两个大规模数据集:SciClaimHunt 和 SciClaimHunt_Num,提出了针对科学论断验证的若干基线模型,并在现有数据集上评估其表现以衡量数据质量。此外,我们进行了人工评估与错误分析,结果表明,SciClaimHunt 和 SciClaimHunt_Num 可作为科学论断验证模型训练的高可靠资源。
原文摘要 · Abstract (English)
Verifying scientific claims presents a significantly greater challenge than verifying political or news-related claims. Unlike the relatively broad audience for political claims, the users of scientific claim verification systems can vary widely, ranging from researchers testing specific hypotheses to everyday users seeking information on a medication. Additionally, the evidence for scientific claims is often highly complex, involving technical terminology and intricate domain-specific concepts that require specialized models for accurate verification. Despite considerable interest from the research community, there is a noticeable lack of large-scale scientific claim verification datasets to benchmark and train effective models. To bridge this gap, we introduce two large-scale datasets, SciClaimHunt and SciClaimHunt_Num, derived from scientific research papers. We propose several baseline models tailored for scientific claim verification to assess the effectiveness of these datasets. Additionally, we evaluate models trained on SciClaimHunt and SciClaimHunt_Num against existing scientific claim verification datasets to gauge their quality and reliability. Furthermore, we conduct human evaluations of the claims in proposed datasets and perform error analysis to assess the effectiveness of the proposed baseline models. Our findings indicate that SciClaimHunt and SciClaimHunt_Num serve as highly reliable resources for training models in scientific claim verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。