从社交媒体推文找科学论断来源,提升溯源准确性
DS@GT at CheckThat! 2025: Exploring Retrieval and Reranking Pipelines for Scientific Claim Source Retrieval on Social Media Discourse
- 设计六种数据增强与七种检索重排组合,优化论文匹配
- 达到MRR@5 0.58,比BM25基线提升0.15
- 适合关注社交媒体科学信息验证的研究者
社交媒体用户常发布无引用来源的科学主张,亟需溯源验证。本文介绍DS@GT团队在CLEF 2025 CheckThat! Lab Task 4b任务中的工作,目标是从推文中隐含的科学主张中检索相关科学论文。团队探索了6种数据增强技术、7种检索与重排管道组合,并微调了一个双编码器模型。最终在任务中取得MRR@5 0.58,排名30支队伍中的第16位,较BM25基线(0.43)提升0.15。代码已开源于GitHub:https://github.com/dsgt-arc/checkthat-2025-swd/tree/main/subtask-4b。
原文摘要 · Abstract (English)
Social media users often make scientific claims without citing where these claims come from, generating a need to verify these claims. This paper details work done by the DS@GT team for CLEF 2025 CheckThat! Lab Task 4b Scientific Claim Source Retrieval which seeks to find relevant scientific papers based on implicit references in tweets. Our team explored 6 different data augmentation techniques, 7 different retrieval and reranking pipelines, and finetuned a bi-encoder. Achieving an MRR@5 of 0.58, our team ranked 16th out of 30 teams for the CLEF 2025 CheckThat! Lab Task 4b, and improvement of 0.15 over the BM25 baseline of 0.43. Our code is available on Github at https://github.com/dsgt-arc/checkthat-2025-swd/tree/main/subtask-4b.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。