arXiv:2507.20700cs.CL2025-07被引 5

小模型在多语言事实核查上胜过大模型,揭示了规模与多样性间的复杂关系。

When Scale Meets Diversity: Evaluating Language Models on Fine-Grained Multilingual Claim Verification

  • 用小模型XLM-R在25种语言上实现更优的细粒度事实验证
  • 小模型宏F1达57.7%,远超最大模型16.9%的表现
  • 适合关注多语言事实核查系统部署的研究者和开发者

多语言虚假信息的快速传播亟需能处理跨语言细粒度真伪判断的自动化验证系统。尽管大语言模型在众多NLP任务中表现优异,其在具有七类细粒度分类的多语言事实核查中的有效性仍研究不足。我们在涵盖25种语言的X-Fact数据集上,评估了五种前沿语言模型:基于编码器的XLM-R(270M参数)和mT5,以及最近的解码器式LLM(Llama 3.1、Qwen 2.5、Mistral Nemo),采用提示和微调两种方法。令人意外的是,XLM-R以57.7%的宏F1显著优于所有测试过的大型模型(最高16.9%),较之前最优结果(41.9%)提升15.8%,建立了新的性能基准。分析发现,大模型存在系统性证据利用困难及在数据不平衡下对高频类别的明显偏倚。这表明,在细粒度多语言事实核查任务中,小型专用模型可能比通用大模型更有效,对事实核查系统的实际部署具有重要启示。

原文摘要 · Abstract (English)

The rapid spread of multilingual misinformation requires robust automated fact verification systems capable of handling fine-grained veracity assessments across diverse languages. While large language models have shown remarkable capabilities across many NLP tasks, their effectiveness for multilingual claim verification with nuanced classification schemes remains understudied. We conduct a comprehensive evaluation of five state-of-the-art language models on the X-Fact dataset, which spans 25 languages with seven distinct veracity categories. Our experiments compare small language models (encoder-based XLM-R and mT5) with recent decoder-only LLMs (Llama 3.1, Qwen 2.5, Mistral Nemo) using both prompting and fine-tuning approaches. Surprisingly, we find that XLM-R (270M parameters) substantially outperforms all tested LLMs (7-12B parameters), achieving 57.7% macro-F1 compared to the best LLM performance of 16.9%. This represents a 15.8% improvement over the previous state-of-the-art (41.9%), establishing new performance benchmarks for multilingual fact verification. Our analysis reveals problematic patterns in LLM behavior, including systematic difficulties in leveraging evidence and pronounced biases toward frequent categories in imbalanced data settings. These findings suggest that for fine-grained multilingual fact verification, smaller specialized models may be more effective than general-purpose large models, with important implications for practical deployment of fact-checking systems.

多语言事实核查小模型胜大模型语言多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。