arXiv:2601.04211cs.CL2026-01

自动评估俄语剧本年龄分级与内容安全,支持快速生成可解释结论。

Qwerty AI: Explainable Automated Age Rating and Content Safety Assessment for Russian-Language Screenplays

  • 用微调的Phi-3-mini模型分段处理剧本,识别五类违规内容。
  • 准确率达80%,分段精度80%-95%,700页剧本2分钟内完成。
  • 无外部调用、低显存限制,适合媒体行业实际生产流程。

我们提出Qwerty AI,一个端到端系统,依据联邦法律436-FZ对俄语剧本进行自动化年龄评级与内容安全评估。系统可处理长达700页的完整剧本,在2分钟内完成,将剧本分割为叙事单元,检测暴力、性内容、脏话、物质滥用、惊吓元素五类违规,并分配0+、6+、12+、16+、18+年龄等级,同时提供可解释的理由。系统采用经过微调的Phi-3-mini模型并使用4比特量化,实现80%的评级准确率和80%-95%的分段精度(受格式影响)。开发受限于:无外部API调用、80GB VRAM上限、平均脚本处理时间<5分钟。部署于雅虎云并启用CUDA加速,展示了其在实际工作流中的应用潜力。该成果在2025年11月的Wink黑客松中取得,解决了俄罗斯传媒业的真实编辑挑战。

原文摘要 · Abstract (English)

We present Qwerty AI, an end-to-end system for automated age-rating and content-safety assessment of Russian-language screenplays according to Federal Law No. 436-FZ. The system processes full-length scripts (up to 700 pages in under 2 minutes), segments them into narrative units, detects content violations across five categories (violence, sexual content, profanity, substances, frightening elements), and assigns age ratings (0+, 6+, 12+, 16+, 18+) with explainable justifications. Our implementation leverages a fine-tuned Phi-3-mini model with 4-bit quantization, achieving 80% rating accuracy and 80-95% segmentation precision (format-dependent). The system was developed under strict constraints: no external API calls, 80GB VRAM limit, and <5 minute processing time for average scripts. Deployed on Yandex Cloud with CUDA acceleration, Qwerty AI demonstrates practical applicability for production workflows. We achieved these results during the Wink hackathon (November 2025), where our solution addressed real editorial challenges in the Russian media industry.

内容安全年龄分级俄语NLP可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。