针对多语言科学溯源,提出基于聚类的难例挖掘方法提升检索精度。
MeVer at CheckThat! 2026: Cluster-Aware Hard-Negative Mining for Multilingual Scientific Source Retrieval
- 利用候选集语义聚类构造更具信息量的难例负样本。
- 局部聚类负样本提升精确率,广义语义负样本增强跨语言覆盖性。
- 结合大模型判别器实现稳定文档选择,适合多语言信息溯源任务。
识别社交媒体中科学主张的来源,需将简短、非正式且多语言的陈述与海量科学文献匹配,其中语义相关论文可能成为训练中的挑战性干扰项或假阴性。本文针对 CheckThat! 2026 多语言科学源检索任务,研究如何在多阶段检索流程中适配难例挖掘策略。提出基于聚类的难例挖掘方法,利用检索候选池的语义结构构建更有效的稠密检索与重排序训练负样本。实验表明,不同难例结构引发不同检索行为:局部聚类负样本有利于提高精确率,而更广泛的非黄金语义负样本则提供更强的候选覆盖,并在跨语言上保持更一致的重排序性能。进一步对比多种基于大模型的证据选择形式,包括直接分类、成对比较和列表重排序提示,发现受限分类提示效果最稳定。最终系统融合稠密检索器、多语言交叉编码器重排序器及选择性大模型分歧解析器,在37个提交中排名第六。结果表明,难例挖掘应作为阶段感知的设计问题,而非单一优化策略。
原文摘要 · Abstract (English)
Identifying the scientific source behind a social media claim requires matching short, informal, and often multilingual claims against large collections of scientific publications, where semantically related papers may act as challenging distractors or false negatives during training. We present our submission to CheckThat! 2026 Task 1 on multilingual scientific-source retrieval, focusing on how hard-negative mining should be adapted to multi-stage retrieval pipelines for scientific source retrieval. We propose cluster-aware hard-negative mining strategies that exploit the semantic structure of retrieved candidate pools in order to construct more informative training negatives for dense retrieval and reranking. Our experiments show that different hard-negative structures induce different retrieval behaviors. Localized cluster negatives tend to favor precision-oriented retrieval, whereas broader non-gold semantic negatives provide stronger candidate coverage and more consistent reranking performance across languages. We further study multiple LLM-based evidence selection formulations, including direct classification, pairwise comparison, and listwise reranking prompts, and find that constrained classification prompts provide the most reliable final document selection. The final system combines a dense retriever, a multilingual cross-encoder reranker, and a selective LLM-based disagreement resolver, ranking 6th among 37 submissions in the shared task evaluation. Overall, our results suggest that hard-negative mining should be treated as a stage-aware design problem rather than as a single retrieval optimization strategy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。