跨语言事实核查中,用多源对齐提升跨语种声明检索准确率。
MultiMind at SemEval-2025 Task 7: Crosslingual Fact-Checked Claim Retrieval via Multi-Source Alignment
- 通过双编码器与对比学习融合多模态源数据,动态学习各来源重要性。
- 在跨语言基准上,检索准确率显著优于基线模型。
- 适合需要多语言事实核查的团队或平台使用。
本文介绍我们为 SemEval-2025 Task 7:多语言与跨语言事实核查声明检索设计的系统。在虚假信息快速传播的时代,高效的事实核查愈发关键。我们提出 TriAligner,一种新颖方法,采用双编码器架构结合对比学习,并融合不同模态下的母语与英文翻译数据。该方法通过学习各来源在对齐中的相对重要性,有效实现多语言声明检索。为增强鲁棒性,我们利用大语言模型进行高效数据预处理与增强,并引入困难负样本以优化表示学习。我们在单语言与跨语言基准上评估该方法,结果表明其在检索准确率和事实核查性能方面均显著优于基线模型。
原文摘要 · Abstract (English)
This paper presents our system for SemEval-2025 Task 7: Multilingual and Crosslingual Fact-Checked Claim Retrieval. In an era where misinformation spreads rapidly, effective fact-checking is increasingly critical. We introduce TriAligner, a novel approach that leverages a dual-encoder architecture with contrastive learning and incorporates both native and English translations across different modalities. Our method effectively retrieves claims across multiple languages by learning the relative importance of different sources in alignment. To enhance robustness, we employ efficient data preprocessing and augmentation using large language models while incorporating hard negative sampling to improve representation learning. We evaluate our approach on monolingual and crosslingual benchmarks, demonstrating significant improvements in retrieval accuracy and fact-checking performance over baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。