系统梳理低资源语言的假信息检测研究,填补多语种检测空白。
Monolingual and Multilingual Misinformation Detection for Low-Resource Languages: A Comprehensive Survey
- 综述低资源语言假信息检测的现有数据集与方法
- 指出数据匮乏与文化语言差异是主要挑战
- 适合关注AI公平性与跨语言治理的研究者
在当今全球数字环境中,假信息已超越语言边界,对内容审核系统构成严峻挑战。当前多数假信息检测方法局限于高资源语言,即少数获得大量研究投入的世界主要语言。本综述全面梳理了低资源语言中单语与多语假信息检测的最新研究,涵盖现有数据集、方法与工具,识别出关键挑战:数据资源不足、模型开发困难、文化与语言背景差异以及实际应用障碍。文章还探讨了新兴技术路径,如语言泛化模型和多模态方法,并强调需改进数据收集实践、推动跨学科合作,加强社会负责任AI研究的激励机制。研究结果凸显了在多元语言与文化背景下应对假信息系统的必要性。
原文摘要 · Abstract (English)
In today's global digital landscape, misinformation transcends linguistic boundaries, posing a significant challenge for moderation systems. Most approaches to misinformation detection are monolingual, focused on high-resource languages, i.e., a handful of world languages that have benefited from substantial research investment. This survey provides a comprehensive overview of the current research on misinformation detection in low-resource languages, both in monolingual and multilingual settings. We review existing datasets, methodologies, and tools used in these domains, identifying key challenges related to: data resources, model development, cultural and linguistic context, and real-world applications. We examine emerging approaches, such as language-generalizable models and multi-modal techniques, and emphasize the need for improved data collection practices, interdisciplinary collaboration, and stronger incentives for socially responsible AI research. Our findings underscore the importance of systems capable of addressing misinformation across diverse linguistic and cultural contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。