检测多轮生成的仇恨视觉故事,发现现有系统常漏判群体性恶意。
Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story Generation

- 构建多轮仇恨故事数据集,测试主流模型生成能力。
- 现有安全系统对群体恶意识别率不足35%,严重漏检。
- 提出交互感知监控与事后分析方法,召回率超90%。
图画书和漫画因易于理解,常被用于传播仇恨叙事,如纳粹宣传画册《毒蘑菇》。近期,像Gemini和GPT-Image这类前沿文本到图像(T2I)系统支持多轮对话生成,能保持角色与场景一致性,使仇恨视觉故事(即传递仇恨叙事的图像序列)变得廉价且可扩展。尽管已有研究关注T2I生成的单张图像中的仇恨内容,但对多图组合构成的群体性恶意仍缺乏探索。本文引入 exttt{HatefulStoryPrompts},包含55个仇恨故事的330个多轮提示,覆盖两种语言和三种视觉风格,评估五款前沿模型在4,950次尝试中的表现:所有模型完成率均超80%,最强达99.0%。进一步在人工标注的 exttt{HatefulVisualStory}数据集(含969组仇恨图像集与990组良性对照)上测试现有审核系统,发现其对群体性恶意识别效果差:专用安全模型最高仅34.9%召回率,强视觉-语言模型达67.5%。最后提出主动防御与事后检测互补方案:交互感知监控在仅提示场景下召回率达97.3%,用户提供首图时为92.6%;事后联合分析已完成图像组的方法召回率可达80.2%。本工作表明,随着图像生成从孤立输出转向连贯视觉叙事,安全机制也需从单图审查升级为对交互状态与图像关系的状态化推理。
原文摘要 · Abstract (English)
Picture books and comics have long been used to disseminate hateful narratives because they are easily understood even by children, as exemplified by the notorious Nazi propaganda picture book \emph{Der Giftpilz}. Recently, frontier text-to-image (T2I) systems such as Gemini and GPT-Image have enabled conversational generation with consistent characters and scenes across turns, making hateful visual stories, namely ordered image groups that collectively convey hateful narratives, cheap and scalable to produce. Although prior work has studied hateful content generation by T2I systems, it focuses on individual images, leaving group-level hateful meaning largely unexplored. We aim to address the gap. Concretely, we introduce \texttt{HatefulStoryPrompts}, comprising 330 multi-turn configurations from 55 hateful stories across two languages and three visual styles, and evaluate five frontier models over 4,950 attempts. Every model completes over 80\% of the stories, with the strongest reaching 99.0\%. We further evaluate existing moderation systems on \texttt{HatefulVisualStory}, a human-labeled dataset of 969 hateful image sets and 990 benign controls, and find that they frequently miss group-level hateful meaning: dedicated safety models achieve at most 34.9\% recall, while a strong vision-language model reaches 67.5\%. Finally, we propose complementary proactive and post-generation defenses. An interaction-aware monitor achieves 97.3\% recall for prompt-only sessions and 92.6\% when the user supplies the first image, while post-generation methods jointly analyzing completed image groups reach 80.2\%. Our work shows that, as image generation evolves from isolated outputs to coherent visual narratives, safety must evolve accordingly, from per-image moderation to stateful reasoning over interactions and image relationships.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。