不依赖训练,用多智能体模拟观众讨论,发现视频争议内容
From Static Analysis to Audience Dissemination: A Training-Free Multimodal Controversy Detection Multi-Agent Framework

- 用三类专业智能体先筛选,分歧时启动观众讨论模拟
- 在有评论和无评论场景下均显著优于现有方法
- 适合社交平台风险管控,无需额外标注数据
多模态争议检测(MCD)旨在识别视频及其用户评论中的争议内容,以支持社交视频平台的风险管理。以往研究将MCD视为静态表征学习任务,直接从视频和评论中提取特征,但未能捕捉不同受众群体的多样视角与评价。受真实内容传播过程启发,我们提出AuDisAgent——一个无需训练的多智能体框架,将MCD重构为动态传播过程。该框架通过结构化多智能体系统显式建模受众传播:首先,视频、评论和交互三类专用筛查智能体分别从视觉、文本和跨模态角度进行初步评估;当三者无法达成一致时,激活观看评议智能体,模拟具有不同背景和立场的受众在筛查后的讨论过程,揭示传播中浮现的潜在争议内容。最后,仲裁智能体基于完整推理链作出最终判断。此外,针对新发布视频缺乏评论的冷启动问题,设计了评论初始化策略,利用语义相似历史视频的公共评论作为初始上下文。在公开数据集上的大量实验表明,该框架在评论丰富与有限两种场景下均显著优于现有最先进方法。
原文摘要 · Abstract (English)
Multimodal controversy detection (MCD) identifies controversial content in videos and their associated user comments, to support risk management for social video platforms.Prior research frames MCD as a static representation learning task, where features are directly extracted from videos and their accompanying comments. However, these methods fail to capture the diverse perspectives and evaluations from different audience groups. Inspired by the real-world process of content dissemination among audiences, we propose AuDisAgent, a training-free multi-agent framework that reformulates MCD as a dynamic propagation process.Our framework explicitly models audience dissemination through a structured multi-agent system. First, three specialized Screening Agents (Video Agent, Comment Agent, and Interaction Agent) conduct initial assessments from visual, textual, and cross-modal perspectives, respectively. For samples where the three agents cannot reach a consensus, a Viewing Panel Agent is activated to simulate post-screening discussions among audiences with diverse backgrounds and stances. This mechanism models how different audience groups interpret and react to the same content, uncovering latent controversial content that may emerge during the dissemination process. Finally, an Arbitration Agent renders the final judgment based on the complete reasoning chain from the preceding steps.In addition, to address the "cold-start" scenario where newly released videos have few or no comments, we design a Comment Bootstrapping Strategy that leverages historical public comments from semantically similar videos as the initial comment context. Extensive experiments on a public dataset demonstrate that our framework significantly outperforms existing state-of-the-art (SOTA) methods in both rich-comment and limited-comment scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。