arXiv:2604.09812cs.CL2026-04

首个多语言事实核查语句嵌入模型,提升跨语言相似性判断与聚类效果。

Claim2Vec: Embedding Fact-Check Claims for Multilingual Similarity and Clustering

论文配图:Claim2Vec: Embedding Fact-Check Claims for Multilingual Similarity and Clustering
图 1 · 摘自论文原文
  • 用对比学习微调多语言编码器,将核查语句映射到语义空间。
  • 在3个数据集上显著提升聚类准确率和嵌入空间结构质量。
  • 支持多语言共现聚类,实现跨语言知识迁移,适合多语种信息核查场景。

重复声明是自动化事实核查系统面临的主要挑战,尤其在多语言环境下。尽管声明匹配和已验证声明检索等任务旨在关联声明对,但如何有效表示一组可通过同一核查结论解决的相似声明(即声明聚类)仍研究不足。为此,我们提出Claim2Vec,首个专为多语言事实核查声明设计的嵌入模型,通过对比学习使用相似多语言声明对微调多语言编码器,将声明映射至优化的语义嵌入空间。在三个数据集、14种多语言嵌入模型和7种聚类算法上的实验表明,Claim2Vec显著提升了聚类性能,尤其在不同聚类配置下均增强了簇标签一致性与嵌入空间几何结构。多语言分析显示,包含多种语言的簇在微调后表现更优,证实了跨语言知识迁移的有效性。

原文摘要 · Abstract (English)

Recurrent claims present a major challenge for automated fact-checking systems designed to combat misinformation, especially in multilingual settings. While tasks such as claim matching and fact-checked claim retrieval aim to address this problem by linking claim pairs, the broader challenge of effectively representing groups of similar claims that can be resolved with the same fact-check via claim clustering remains relatively underexplored. To address this gap, we introduce Claim2Vec, the first multilingual embedding model optimized to represent fact-check claims as vectors in an improved semantic embedding space. We fine-tune a multilingual encoder using contrastive learning with similar multilingual claim pairs. Experiments on the claim clustering task using three datasets, 14 multilingual embedding models, and 7 clustering algorithms demonstrate that Claim2Vec significantly improves clustering performance. Specifically, it enhances both cluster label alignment and the geometric structure of the embedding space across different cluster configurations. Our multilingual analysis shows that clusters containing multiple languages benefit from fine-tuning, demonstrating cross-lingual knowledge transfer.

多语言嵌入模型事实核查聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。