为视频平台内容审核模型提供故障诊断与干预方法
From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation

- 构建可诊断的失败分类体系,定位审核模型失效原因
- 在真实全球视频流上支持超100个区域,应对噪声与多样性挑战
- 帮助平台从盲目调参转向精准修复,提升安全模型开发效率
大规模视频与直播平台的内容审核对模型提出严苛要求:需适配平台特有数据分布、政策目标及产品级安全约束。现有通用预训练模型或外部API难以满足这些需求,平台必须自研模型,但现有研究多聚焦架构设计与基准性能,缺乏对故障定位与针对性改进的系统指导。实际部署中的失败往往成因复杂,同类问题可能源于不同机制,若无精准干预,优化只能依赖经验试错,导致性能提升难归因、问题难溯源。为此,本文提出一套面向工业级音视频语言模型(AVLM)开发的诊断方法:将模型失效映射为可观测的失败模式分类,并关联对应的干预策略。该方法在某大型视频直播平台的AVLM全生命周期中落地应用,支持超过100个地区,处理来自全球流量的噪声大、语义模糊且高度多样化的视听内容。
原文摘要 · Abstract (English)
Industry-scale video and live-streaming moderation imposes requirements that are difficult to satisfy with generic pretrained public models or external APIs, including adaptation to platform-specific data distributions, policy-specific objectives, and product-level safety constraints. As a result, platforms must undertake internal model development, naturally turning to shared public research for guidance. However, existing multimodal foundation-model studies primarily report architectures, training recipes, data scaling strategies, and benchmark results, but provide less systematic guidance on how failures should be localized and translated into targeted model-development interventions. Interventions are essential because deployment failures are rarely self-explanatory. Similar failures can originate from different causes. Without targeted interventions, improvement reduces to heuristic trial-and-error, where benchmark improvements are weakly attributable, and failures are difficult to trace to their underlying causes. To address this gap, we present a diagnostic methodology for industry-scale Audio-Visual-Language Models AVLM development. The methodology maps model failures into a taxonomy of observable failure signatures and links each class of failure to an intervention space. We instantiate this methodology across the development and alignment lifecycle of an AVLM foundation model for a large-scale video and live-streaming platform. The resulting system supports over 100 regions and is designed for noisy, ambiguous, and highly diverse content drawn from global platform traffic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。