arXiv:2504.16627cs.CL2025-04ACL被引 1

用翻译提升多语言事实核查召回,能在消费级显卡运行

TIFIN India at SemEval-2025: Harnessing Translation to Overcome Multilingual IR Challenges in Fact-Checked Claim Retrieval

  • 先用微调嵌入模型检索,再用大模型重排
  • 单语和跨语言测试集准确率分别达0.938和0.81025
  • 全程可复现且仅需普通显卡,适合实际部署

我们解决在单语和跨语言场景下检索已核实声明的挑战,这对应对全球虚假信息传播至关重要。采用两阶段策略:基于微调嵌入模型的可靠基线检索系统,以及基于大语言模型的重排器。关键贡献在于证明大模型翻译能有效克服多语言信息检索障碍。同时,确保整个流程可在消费级显卡上复现。最终集成系统在单语和跨语言测试集上的success@10得分分别为0.938和0.81025。

原文摘要 · Abstract (English)

We address the challenge of retrieving previously fact-checked claims in monolingual and crosslingual settings - a critical task given the global prevalence of disinformation. Our approach follows a two-stage strategy: a reliable baseline retrieval system using a fine-tuned embedding model and an LLM-based reranker. Our key contribution is demonstrating how LLM-based translation can overcome the hurdles of multilingual information retrieval. Additionally, we focus on ensuring that the bulk of the pipeline can be replicated on a consumer GPU. Our final integrated system achieved a success@10 score of 0.938 and 0.81025 on the monolingual and crosslingual test sets, respectively.

多语言检索事实核查大模型应用轻量化部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。