arXiv:2505.15063cs.CL2025-05EMNLP被引 5

首个针对乌尔都语的自动事实核查框架,解决低资源语言可信度问题。

UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking

  • 构建双基准:声明验证与问答事实性评估,专为乌尔都语设计
  • 翻译增强策略使大模型表现优于纯乌尔都语检索,提升准确率
  • 开源数据与代码,适合低资源语言研究者与NLP工程师使用

大型语言模型在乌尔都语等低资源语言中的事实可靠性引发关注。现有系统多面向英语,难以覆盖全球超2亿乌尔都语使用者。本文提出首个乌尔都语事实核查基准——UrduFactBench(用于声明验证)与UrduFactQA(用于问答事实性评估),均由母语者参与多阶段标注完成。为此构建了UrduFactCheck框架,结合单语与翻译式证据检索策略,缓解高质量乌尔都语证据稀缺问题。在12个大模型上评估显示,翻译增强管道性能显著优于纯单语方案。结果揭示开源模型在乌尔都语上的持续挑战,凸显专用资源的重要性。所有代码与数据已公开于https://github.com/mbzuai-nlp/UrduFactCheck。

原文摘要 · Abstract (English)

The rapid adoption of Large Language Models (LLMs) has raised important concerns about the factual reliability of their outputs, particularly in low-resource languages such as Urdu. Existing automated fact-checking systems are predominantly developed for English, leaving a significant gap for the more than 200 million Urdu speakers worldwide. In this work, we present UrduFactBench and UrduFactQA, two novel hand-annotated benchmarks designed to enable fact-checking and factual consistency evaluation in Urdu. While UrduFactBench focuses on claim verification, UrduFactQA targets the factuality of LLMs in question answering. These resources, the first of their kind for Urdu, were developed through a multi-stage annotation process involving native Urdu speakers. To complement these benchmarks, we introduce UrduFactCheck, a modular fact-checking framework that incorporates both monolingual and translation-based evidence retrieval strategies to mitigate the scarcity of high-quality Urdu evidence. Leveraging these resources, we conduct an extensive evaluation of twelve LLMs and demonstrate that translation-augmented pipelines consistently enhance performance compared to monolingual ones. Our findings reveal persistent challenges for open-source LLMs in Urdu and underscore the importance of developing targeted resources. All code and data are publicly available at https://github.com/mbzuai-nlp/UrduFactCheck.

事实核查低资源语言大模型评估乌尔都语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。