arXiv:2601.16027cs.AI2026-01被引 1

用跨直播会话证据提升风险识别能力,兼顾实时性与可解释性。

Deja Vu in Plots: Leveraging Cross-Session Evidence with Retrieval-Augmented LLMs for Live Streaming Risk Assessment

  • 通过检索增强的LLM指导小模型,挖掘跨会话行为模式
  • 在大规模工业数据集上达到顶尖性能,支持实时部署
  • 输出可解释信号,适合平台内容审核场景

直播兴起推动了线上互动的爆发式增长,但也带来了诈骗和协同攻击等复杂风险。由于有害行为常呈渐进式累积且跨流重复出现,传统检测方法效果有限。为此,我们提出CS-VAR(跨会话证据感知的检索增强检测器),通过轻量级领域专用模型实现快速会话级风险判断。该模型在训练中由大型语言模型(LLM)引导,基于检索到的跨会话行为证据进行推理,并将局部到全局的洞察转移给小模型。这一设计使小模型能识别跨流重复模式,实现结构化风险评估,同时保持实时部署效率。在大规模工业数据集上的离线实验及在线验证均表明,CS-VAR表现优于现有方法。此外,系统提供可解释的局部信号,有效支持真实场景下的直播内容治理。

原文摘要 · Abstract (English)

The rise of live streaming has transformed online interaction, enabling massive real-time engagement but also exposing platforms to complex risks such as scams and coordinated malicious behaviors. Detecting these risks is challenging because harmful actions often accumulate gradually and recur across seemingly unrelated streams. To address this, we propose CS-VAR (Cross-Session Evidence-Aware Retrieval-Augmented Detector) for live streaming risk assessment. In CS-VAR, a lightweight, domain-specific model performs fast session-level risk inference, guided during training by a Large Language Model (LLM) that reasons over retrieved cross-session behavioral evidence and transfers its local-to-global insights to the small model. This design enables the small model to recognize recurring patterns across streams, perform structured risk assessment, and maintain efficiency for real-time deployment. Extensive offline experiments on large-scale industrial datasets, combined with online validation, demonstrate the state-of-the-art performance of CS-VAR. Furthermore, CS-VAR provides interpretable, localized signals that effectively empower real-world moderation for live streaming.

风险检测直播安全检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。