多语言多智能体模型可识别跨语言翻译、题型变换等新型虚假信息攻击。
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
- 构建多语言多智能体系统,结合检索增强生成检测多种伪装手法。
- 在英法西阿印中六种语言上验证对换语、长句膨胀等攻击的检测效果。
- 支持网页插件部署,适合平台方快速集成到真实网络环境使用。
数字平台中虚假信息的快速传播威胁公共讨论、情绪稳定与决策质量。尽管已有研究探讨了虚假信息检测中的各类对抗攻击,但本文首次系统研究了英语、法语、西班牙语、阿拉伯语、印地语和中文之间的语言切换后翻译、查询长度膨胀后摘要化,以及结构重排为选择题等攻击手段。本文提出一种多语言多智能体大语言模型框架,采用检索增强生成技术,可作为网页插件部署于在线平台。研究强调了人工智能在抵御多样化攻击、维护网络事实真实性方面的重要性,同时展示了插件式部署在实际应用中的可行性。
原文摘要 · Abstract (English)
The rapid spread of misinformation on digital platforms threatens public discourse, emotional stability, and decision-making. While prior work has explored various adversarial attacks in misinformation detection, the specific transformations examined in this paper have not been systematically studied. In particular, we investigate language-switching across English, French, Spanish, Arabic, Hindi, and Chinese, followed by translation. We also study query length inflation preceding summarization and structural reformatting into multiple-choice questions. In this paper, we present a multilingual, multi-agent large language model framework with retrieval-augmented generation that can be deployed as a web plugin into online platforms. Our work underscores the importance of AI-driven misinformation detection in safeguarding online factual integrity against diverse attacks, while showcasing the feasibility of plugin-based deployment for real-world web applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。