用大模型跨语言检测已核查过的内容,提升全球辟谣效率
Large Language Models for Multilingual Previously Fact-Checked Claim Detection
- 评测7个大模型在20种语言上识别已核查声明的能力
- 高资源语言表现良好,低资源语言通过英译提升效果
- 适合从事多语言信息治理与反虚假信息研究者参考
在虚假信息泛滥的时代,人类核实员常需重复验证已被他国或他语种处理过的声明。由于虚假信息跨越语言边界,自动识别跨语言已核查声明成为关键任务。本文首次全面评估大语言模型(LLMs)在多语言已核查声明检测中的表现。我们在20种语言中测试了7个LLMs,在单语言和跨语言设置下进行评估。结果显示,尽管高资源语言表现优异,但低资源语言仍面临挑战;将原文翻译为英语对低资源语言有显著帮助。这些发现凸显了LLMs在多语言已核查声明检测中的潜力,并为该方向的后续研究奠定基础。
原文摘要 · Abstract (English)
In our era of widespread false information, human fact-checkers often face the challenge of duplicating efforts when verifying claims that may have already been addressed in other countries or languages. As false information transcends linguistic boundaries, the ability to automatically detect previously fact-checked claims across languages has become an increasingly important task. This paper presents the first comprehensive evaluation of large language models (LLMs) for multilingual previously fact-checked claim detection. We assess seven LLMs across 20 languages in both monolingual and cross-lingual settings. Our results show that while LLMs perform well for high-resource languages, they struggle with low-resource languages. Moreover, translating original texts into English proved to be beneficial for low-resource languages. These findings highlight the potential of LLMs for multilingual previously fact-checked claim detection and provide a foundation for further research on this promising application of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。