通过自适应切分文章,提升多语言事实核查的证据完整性和准确性
SEEK: Semantic Evidence Extraction via Adaptive ChunKing for Multilingual Fact-Checking

- 根据语义转折点自动切割文章,生成连贯证据块
- 在多语言数据集上将宏观F1提升最高达20%
- 适合需要高质量跨语言事实验证的研究与应用
多语言事实核查需要既相关又完整的证据以实现可靠的事实判断。现有系统常依赖搜索片段、句子级证据或局部段落,易遗漏关键上下文,导致证据碎片化。为此,我们提出SEEK——一种基于自适应切分的语义证据提取框架,通过识别语义主题转换,从全文中构建连贯的证据块,并保留局部验证上下文。这些块由多语言编码器处理,再使用LoRA适配器微调多语言大模型进行真伪判断。在X-FACT和RU22Fact数据集上的实验表明,SEEK相较语义分块提升宏平均F1达10%,相较句子分块提升19%,相较搜索片段基线提升20%。证据完整性与重要性分析显示,SEEK能更好保留验证上下文,支持更可靠的多语言事实核查。
原文摘要 · Abstract (English)
Multilingual fact verification requires evidence that is both relevant and sufficiently complete for reliable factuality prediction. However, existing systems often rely on search snippets, sentence-level evidence, or locally segmented passages, which can miss decisive context and produce fragmented evidence. To overcome these limitations, we propose SEEK, a Semantic Evidence Extraction with an adaptive chunKing framework that constructs coherent evidence chunks from full fact-checking articles by identifying semantic topic transitions and preserving local verification context. The constructed chunks are encoded using a multilingual encoder and then multilingual LLMs are finetuned using LoRA adapter for veracity prediction. Experiments on X-FACT and RU22Fact show that SEEK improves macro-f1 by up to 10% over semantic chunking, 19% over sentence chunking, and 20% over search-snippet baselines. Evidence completeness and significance analyses further show that SEEK preserves richer verification context and enables more reliable multilingual fact-checking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。