arXiv:2604.17020cs.CLcs.AI2026-04ACL

用角色模拟生成多样有害内容,提升检测系统评测可靠性

Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation

论文配图:Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation
图 1 · 摘自论文原文
  • 构建双维度角色画像,融合身份、兴趣与恶意策略
  • 合成内容在危害性、挑战性和多样性上均超越现有基准
  • 适合用于测试和优化内容安全模型的鲁棒性

静态基准在可扩展性和多样性方面存在局限,且可能受大规模预训练语料污染。为此,我们提出一种基于角色引导的大型语言模型代理框架,用于合成有害内容。通过整合人口属性、主题兴趣与情境化恶意策略,构建二维用户角色,实现多样化且上下文相关的有害交互模拟。从危害性、挑战度和多样性三个维度评估,人工与模型双重验证表明,该框架具有高生成成功率。多系统实验显示,合成场景比现有基准更难检测。多角度分析证实,其语言和主题多样性可媲美人工标注数据集,证明该框架是评测有害内容检测系统鲁棒性的有效工具。

原文摘要 · Abstract (English)

Static benchmarks for harmful content detection face limitations in scalability and diversity, and may also be affected by contamination from web-scale pre-training corpora. To address these issues, we propose a framework for synthesizing harmful content, leveraging persona-guided large language model (LLM) agents. Our approach constructs two-dimensional user personas by integrating demographic identities and topical interests with situational harmful strategies, enabling the simulation of diverse and contextually grounded harmful interactions. We evaluate the framework along three dimensions: harmfulness, challenge level, and diversity. Both human and LLM-based evaluations confirm that our framework achieves a high harmful generation success rate. Experiments across multiple detection systems reveal that our synthetic scenarios are more challenging to detect than those in existing benchmarks. Furthermore, a multi-faceted analysis confirms that our approach achieves linguistic and topical diversity comparable to human-curated datasets, establishing our framework as an effective tool for robust stress-testing of harmful content detection systems.

有害内容检测角色模拟评测框架LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。