arXiv:2607.15875cs.IR2026-07中稿 · CLEF 2026

提升科学声明溯源准确率,通过风格迁移与重排序优化检索效果。

Scientific Claim-Source Retrieval Revisited: A Comparative Study of Style Transfer and Re-Ranking

论文配图:Scientific Claim-Source Retrieval Revisited: A Comparative Study of Style Transfer and Re-Ranking
图 1 · 摘自论文原文
  • 将中文声明翻译成英文,显著优于原语种和双语表示。
  • 引入出版物元数据,提升检索性能;验证型重排序达MRR@5 0.758最佳。
  • 适合需要高精度科学信息溯源的研究者与平台开发者。

社交媒体上的科学声明常难以验证,易引发虚假信息传播。为应对这一挑战,自动化事实核查系统需依赖科学声明-来源检索任务,即识别给定声明所对应的原始文献。然而,声明与原文在语言、风格和具体性上差异较大,导致检索困难。本文在CheckThat! 2026基准上对比了稀疏与稠密检索模型的表现。结果表明,将声明翻译为英文优于原始语种及双语表示;引入出版物元数据可进一步提升检索效果,捕捉间接来源引用。我们分析了四种风格迁移方法,发现其普遍提升多数模型的检索性能,但最优风格取决于底层检索目标。此外,我们提出三种基于归属、实体重叠和验证推理的新型重排序模型。其中,验证型重排序在语义相似性基础上取得额外增益,整体表现最佳,MRR@5达到0.758。

原文摘要 · Abstract (English)

Scientific claims shared on social media are often difficult to verify and may contribute to the spread of misinformation. To address this challenge, automated fact verification systems require scientific claim-source retrieval, the task of identifying the source publication underlying a given claim. However, claims often differ substantially from their source publications in language, style, and specificity, making retrieval challenging. We present a comparative study of scientific claim-source retrieval on the CheckThat! 2026 benchmark across sparse and dense retrieval models. Our results show that translating claims into English outperforms both original and bilingual claim representations, while incorporating publication metadata provides additional retrieval gains by capturing indirect source references. In addition, we analyze four style transfer approaches and find that they improve retrieval performance for most models, although the optimal style depends on the underlying retrieval objective. Finally, we investigate similarity- and signal-based re-ranking approaches, introducing three novel re-ranking models based on attribution, entity overlap, and verification-based reasoning. Verification-based re-ranking yields additional gains beyond semantic similarity and achieves the best overall performance with an MRR@5 of 0.758.

信息溯源风格迁移重排序科学验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。