游戏聊天毒性检测需结合上下文与领域知识,提升识别准确率。
Context-Aware Toxicity Detection in Multiplayer Games: Integrating Domain-Adaptive Pretraining and Match Metadata
- 用领域自适应预训练增强模型对游戏俚语的理解
- 融合比赛元数据和历史对话,性能显著优于孤立消息检测
- 适合游戏厂商构建智能内容审核系统
竞技类在线视频游戏中毒性言论的危害广受关注,但因其依赖上下文、常跨越多条消息或受非文本互动影响,传统检测方法难以应对。尤其在游戏场景中,玩家使用大量专有缩写、俚语和拼写错误,且毒性行为稀少,标准模型难以有效识别。本文将RoBERTa大模型适配至游戏领域,通过引入比赛元数据和历史交互信息,增强预训练嵌入表示,以捕捉玩家互动的深层语义。基于DOTA 2与Call of Duty®: Modern Warfare® III两个游戏数据集,实验证明:结合元数据与先前对话能显著提升检测效果,明确了上下文与领域适配的关键作用。本研究强调了上下文感知与领域定制化在主动内容治理中的重要性。
原文摘要 · Abstract (English)
The detrimental effects of toxicity in competitive online video games are widely acknowledged, prompting publishers to monitor player chat conversations. This is challenging due to the context-dependent nature of toxicity, often spread across multiple messages or informed by non-textual interactions. Traditional toxicity detectors focus on isolated messages, missing the broader context needed for accurate moderation. This is especially problematic in video games, where interactions involve specialized slang, abbreviations, and typos, making it difficult for standard models to detect toxicity, especially given its rarity. We adapted RoBERTa LLM to support moderation tailored to video games, integrating both textual and non-textual context. By enhancing pretrained embeddings with metadata and addressing the unique slang and language quirks through domain adaptive pretraining, our method better captures the nuances of player interactions. Using two gaming datasets - from Defense of the Ancients 2 (DOTA 2) and Call of Duty$^\circledR$: Modern Warfare$^\circledR$III (MWIII) we demonstrate which sources of context (metadata, prior interactions...) are most useful, how to best leverage them to boost performance, and the conditions conducive to doing so. This work underscores the importance of context-aware and domain-specific approaches for proactive moderation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。