arXiv:2605.31136cs.CL2026-05ACL

构建多语言维基引用检测数据集,证明小模型比大模型更适合低资源语言事实核查

Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages

论文配图:Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages
图 1 · 摘自论文原文
  • 构建覆盖18种语言的多语言引用检测数据集,分三类资源等级
  • 小模型经编码器目标微调后性能超越提示大模型,跨语言迁移效果佳
  • 适合低资源社区使用,为维基百科内容可信度提升提供轻量级方案

在自动化事实核查中,检查必要性检测依据领域标准识别需验证的陈述。在维基百科中,该任务表现为引用必要性检测(CND),即标记缺乏支持引文的陈述。然而现有研究大多忽视低资源语言,且近期事实核查流程依赖大型语言模型(LLMs),这对低资源组织不可及。本文提出MCN,一个涵盖18种语言、三类资源水平的多语言CND语料库,并对小型解码器语言模型(SLMs)进行了广泛研究。实验表明,使用编码器式目标微调的SLMs在所有语言上显著优于提示式LLMs。我们还开展了首个跨语言CND研究,发现仅用英语数据微调的SLMs在多数目标语言上表现优于需大量适配的LLMs。研究结果对低资源维基社区具有重要意义,表明针对任务的紧凑专用模型优于通用大模型。所有数据与代码已公开于https://github.com/gerritq/mcn。

原文摘要 · Abstract (English)

In automated fact-checking (AFC), check-worthiness detection identifies claims requiring verification based on domain-specific criteria. On Wikipedia, this task instantiates as Citation Needed Detection (CND), which flags claims lacking supporting citations. However, existing research has largely overlooked lower-resource languages, and recent AFC pipelines rely on large language models (LLMs), which are inaccessible to low-resource organizations. We introduce MCN, a multilingual CND corpus spanning 18 languages across three resource levels, on which we conduct an extensive study of small decoder-based language models (SLMs). Our experiments show that SLMs fine-tuned with an encoder-style objective substantially outperform prompted LLMs across languages. We further present one of the first studies on cross-lingual CND, demonstrating that SLMs fine-tuned solely on English claims surpass LLMs, even with little to no target-language adaptation. Our findings have important implications for lower-resource Wikipedia communities and suggest that compact, task-specific models are preferable to LLMs for CND. We release all data and code at https://github.com/gerritq/mcn

事实核查多语言低资源维基百科

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。