arXiv:2605.12856cs.AIcs.SI2026-05

通过多轮对话识别恶意代理的真实意图,提升系统安全。

Moltbook Moderation: Uncovering Hidden Intent Through Multi-Turn Dialogue

论文配图:Moltbook Moderation: Uncovering Hidden Intent Through Multi-Turn Dialogue
图 1 · 摘自论文原文
  • 用多轮对话和概率采样探查代理隐藏意图
  • 在多种对抗设置下准确识别恶意行为,误报率低
  • 适合开放环境中需要深度意图检测的平台

多智能体系统的兴起带来了超越内容过滤的新监管挑战。具有恶意意图的代理可能输出看似无害的内容以规避基于内容的审核,同时通过整体交互模式实施破坏性行为。为此,我们提出BOT-MOD(BOT-MODeration)框架,将检测基础从内容层面转向代理意图。该方法通过吉布斯采样对候选意图假设进行多轮交互式推演,逐步缩小合理目标空间,从而识别深层行为动机。为评估该方法,我们基于Moltbook构建了数据集,涵盖真实社区结构下的多样良性与恶意行为。实验表明,BOT-MOD在多种对抗配置下能可靠识别代理意图,且对良性行为保持低误报率。本工作为开放多智能体环境中的可扩展、意图感知式监管奠定了基础。

原文摘要 · Abstract (English)

The emergence of multi-agent systems introduces novel moderation challenges that extend beyond content filtering. Agents with malicious intent may contribute harmful content that appears benign to evade content-based moderation, while compromising the system through exploitative and malicious behavior manifested across their overall interaction patterns within the community. To address this, we introduce BOT-MOD (BOT-MODeration), a moderation framework that grounds detection in agent intent rather than traditional content level signals. BOT-MOD identifies the underlying intent by engaging with the target agent in a multi-turn exchange guided by Gibbs-based sampling over candidate intent hypotheses. This progressively narrows the space of plausible agent objectives to identify the underlying behavior. To evaluate our approach, we construct a dataset derived from Moltbook that encompasses diverse benign and malicious behaviors based on actual community structures, posts, and comments. Results demonstrate that BOT-MOD reliably identifies agent intent across a range of adversarial configurations, while maintaining a low false positive rate on benign behaviors. This work advances the foundation for scalable, intent-aware moderation of agents in open multi-agent environments.

智能体安全意图识别多轮对话内容审核

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。