用多智能体辩论机制提升内容安全审查的透明度与隐性风险识别能力
Aetheria: A multimodal interpretable content safety framework based on multi-agent debate and collaboration
- 五智能体协同辩论,结合检索增强生成实现多模态内容深度分析
- 在AIR-Bench上显著提升隐性风险识别准确率,审计报告可追溯
- 适合需要可解释AI审核的平台方、政策制定者及可信AI研究者
数字内容的指数增长给内容安全带来严峻挑战。现有审核系统多基于单模型或固定流程,难以识别隐性风险且判断过程不透明。为此,我们提出Aetheria,一种基于多智能体辩论与协作的多模态可解释内容安全框架。该框架采用五个核心智能体的协同架构,通过动态互辩机制对多模态内容进行深度分析与裁决,并依托RAG知识检索提供支撑。在自建基准AIR-Bench上的全面实验表明,Aetheria不仅能生成详细可追溯的审计报告,且在整体内容安全准确性上显著优于基线模型,尤其在隐性风险识别方面表现突出。该框架建立了透明可解释的新范式,大幅推动可信AI内容审核的发展。
原文摘要 · Abstract (English)
The exponential growth of digital content presents significant challenges for content safety. Current moderation systems, often based on single models or fixed pipelines, exhibit limitations in identifying implicit risks and providing interpretable judgment processes. To address these issues, we propose Aetheria, a multimodal interpretable content safety framework based on multi-agent debate and collaboration.Employing a collaborative architecture of five core agents, Aetheria conducts in-depth analysis and adjudication of multimodal content through a dynamic, mutually persuasive debate mechanism, which is grounded by RAG-based knowledge retrieval.Comprehensive experiments on our proposed benchmark (AIR-Bench) validate that Aetheria not only generates detailed and traceable audit reports but also demonstrates significant advantages over baselines in overall content safety accuracy, especially in the identification of implicit risks. This framework establishes a transparent and interpretable paradigm, significantly advancing the field of trustworthy AI content moderation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。