系统梳理网络毒害行为的成因、检测与治理方法,为AI时代构建安全对话环境提供指南。
Toxicity in Online Platforms and AI Systems: A Survey of Needs, Challenges, Mitigations, and Future Directions
- 从多角度构建毒性分类体系,涵盖语言、图像、视频等多模态内容
- 总结主流英文语料库与大模型在毒性检测中的表现差异
- 指出当前研究在可解释性、自适应性和评估标准上的不足
数字通信系统和在线平台的设计无意中助长了隐性传播毒害行为的现象,导致对毒害内容的被动反应。网络内容及人工智能系统中的毒害行为已成为全球个体与集体福祉的重大挑战,其危害远超普遍认知。毒害表现为语言、图像和视频等形式,其含义依赖具体语境。因此,建立全面的毒性分类体系对于主动识别和缓解毒害至关重要。深入理解毒害现象有助于设计切实可行的检测与治理方案。现有文献的分类仅关注该复杂问题的有限方面,且多采取被动应对策略。本综述旨在从多个视角构建毒性综合分类体系,通过理解人工智能时代的社会背景与环境,提出整体性解释框架。文章总结了相关数据集及针对大语言模型、社交媒体平台等在线平台的毒害检测与缓解研究,聚焦英文文本模式。最后,基于数据集、缓解策略、大模型特性、可适应性、可解释性与评估体系等方面,指出了当前研究的空白。
原文摘要 · Abstract (English)
The evolution of digital communication systems and the designs of online platforms have inadvertently facilitated the subconscious propagation of toxic behavior. Giving rise to reactive responses to toxic behavior. Toxicity in online content and Artificial Intelligence Systems has become a serious challenge to individual and collective well-being around the world. It is more detrimental to society than we realize. Toxicity, expressed in language, image, and video, can be interpreted in various ways depending on the context of usage. Therefore, a comprehensive taxonomy is crucial to detect and mitigate toxicity in online content, Artificial Intelligence systems, and/or Large Language Models in a proactive manner. A comprehensive understanding of toxicity is likely to facilitate the design of practical solutions for toxicity detection and mitigation. The classification in published literature has focused on only a limited number of aspects of this very complex issue, with a pattern of reactive strategies in response to toxicity. This survey attempts to generate a comprehensive taxonomy of toxicity from various perspectives. It presents a holistic approach to explain the toxicity by understanding the context and environment that society is facing in the Artificial Intelligence era. This survey summarizes the toxicity-related datasets and research on toxicity detection and mitigation for Large Language Models, social media platforms, and other online platforms, detailing their attributes in textual mode, focused on the English language. Finally, we suggest the research gaps in toxicity mitigation based on datasets, mitigation strategies, Large Language Models, adaptability, explainability, and evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。