用强化学习优化游戏毒性检测,提升效率与准确性。
Reinforcement Learning for Efficient Toxicity Detection in Competitive Online Video Games
- 基于领域知识设计上下文赌博算法,动态决策监测策略。
- 在《使命召唤:现代战争3》数据上优于仅依赖历史行为的基线方法。
- 适合游戏平台安全团队部署,可快速落地提升监管效能。
在线平台需主动检测并应对不良行为,以集中资源于高风险场景。本文研究竞技类在线视频游戏中毒性行为的高效采样检测问题。为做出最优监控决策,服务运营商需预估毒性行为的发生概率;若缺乏预测模型,则需实时构建。为此,我们提出一种基于领域专家判断的上下文赌博算法,利用少量相关变量进行监控决策,平衡探索与利用,优化长期效果,并专为生产环境部署设计。基于《使命召唤:现代战争3》的真实数据,实验表明该算法持续优于仅依赖玩家历史行为的基线方法。这一结果揭示了毒性行为的本质特征,也展示了如何将领域知识融入系统以识别和缓解毒性,最终营造更安全、愉快的游戏体验。
原文摘要 · Abstract (English)
Online platforms take proactive measures to detect and address undesirable behavior, aiming to focus these resource-intensive efforts where such behavior is most prevalent. This article considers the problem of efficient sampling for toxicity detection in competitive online video games. To make optimal monitoring decisions, video game service operators need estimates of the likelihood of toxic behavior. If no model is available for these predictions, one must be estimated in real time. To close this gap, we propose a contextual bandit algorithm that makes monitoring decisions based on a small set of variables that, according to domain expertise, are associated with toxic behavior. This algorithm balances exploration and exploitation to optimize long-term outcomes and is deliberately designed for easy deployment in production. Using data from the popular first-person action game Call of Duty: Modern Warfare III, we show that our algorithm consistently outperforms baseline algorithms that rely solely on players' past behavior. This finding has substantive implications for the nature of toxicity. It also illustrates how domain expertise can be harnessed to help video game service operators identify and mitigate toxicity, ultimately fostering a safer and more enjoyable gaming experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。